llms

99 stories and discussions about llms, aggregated from every source we track.

1.

Are you steering towards AI burnout? Afraid of loosing your job to someone with little programming skills, no aspirations to quality, and a huge Claude account? Disappointed about the code quality in your projects, or…

330 points•signa11•5 days ago•337 comments•
2.

Run Q4 small language models entirely in the browser with WebGPU. Chat, benchmark, and compare PetitGPT, SmolLM2, and friends on-device.

278 points•logicallee•2 days ago•113 comments•
3.

AI drastically shortened its design time; it will only get faster

201 points•maxall4•12 days ago•136 comments•
4.

I was intrigued by Jev and the self-hostable projects appearing around it, such as OpenJev and SemIf . Reading about them introduced me ...

148 points•allanrbo•5 days ago•44 comments•
6.

Are you steering towards AI burnout? Afraid of loosing your job to someone with little programming skills, no aspirations to quality, and a huge Claude account? Disappointed about the code quality in your projects, or…

28 points•eatonphil•5 days ago•2 comments
7.
16 points•leonardovb•12 days ago•26 comments•
8.

You can't get them to stop talking about spines!

13 points•zdw•11 days ago•2 comments•
9.

What the independent measurements of TypeSafe's Jev found in its first eight days: arXiv preprints, GitHub evaluations and blog benchmarks, each traced to its primary source. Accuracy, calibration, speed, cost, failure modes, the prior art, the open alternatives, and what is still unmeasured.

9 points•xbill•7 days ago•1 comment
10.

An esoteric programming language built to break LLMs and coding agents.

9 points•dom96•11 days ago•2 comments•
11.
6 points•vrnithinkumar•3 days ago•2 comments•
12.
6 points•avestura•3 days ago•0 comments•
13.

This is a submission for the Kaggle Benchmarking Challenge What I Benchmarked ...

6 points•nishikantaray•3 days ago•5 comments
14.
6 points•jwbobbink•6 days ago•0 comments•
15.
6 points•ovyan•9 days ago•2 comments•
16.

The framework for programming—rather than prompting—language models.

5 points•mpweiher•3 days ago•1 comment•
17.

Every Apple Silicon Mac can run AI locally. Here's what model sizes actually fit your machine, the tools to use, and the daily-driver use nobody talks about.

5 points•donk8r•6 days ago•0 comments•
18.
4 points•cpa•6 days ago•1 comment•
19.
4 points•Brajeshwar•8 days ago•1 comment•
20.

A from-scratch derivation of REINFORCE for language models

3 points•gmays•about 11 hours ago•0 comments•
21.

Plus Claude Opus 5.5, GPT-6 Sol, GPT-6 Luna, and a new price war

3 points•swolpers•2 days ago•0 comments•
22.
3 points•bruhfeefee•3 days ago•0 comments•
23.

StarSkirmish: where LLMs write code to play StarCraft.

3 points•__cayenne__•4 days ago•1 comment•
24.

"Millions of people around the world will soon consider large models ‘hoovering up’ all their work to be an astonishing theft of unprecedented proportions."

3 points•xqcgrek2•6 days ago•1 comment•
25.

Using "think" to describe what an LLM does reminds me of the 16th century, when astronomers didn't really believe in the heliocentric model, but used it for calculations anyway because the math was simpler. We'll start…

3 points•fooker•6 days ago•3 comments•
26.

Benchmark of TypeSafe's Jev against Sonnet 5, GPT-5 nano and local LLMs on 770 Reddit AITA verdicts: Brier scores, latency and cost - dchristopoulos/jev-aita

3 points•dchristopoulos•7 days ago•0 comments•
27.
3 points•Dota-12•8 days ago•0 comments•
28.
3 points•vietthangif•9 days ago•0 comments•
29.

The model deciding whether to wake someone up doesn't need to sing Bohemian Rhapsody in Klingon. Working with Jev has me thinking about how much less we should ask of AI.

3 points•pgr0ss•9 days ago•0 comments•
30.

Some miscellaneous thoughts on what LLMs are changing in my field.

3 points•luu•11 days ago•0 comments•
31.

AI drastically shortened its design time; it will only get faster

3 points•vaultdmn•12 days ago•0 comments
32.

VRAM for LLMs is a bandwidth problem: every token streams the whole model from memory. Bandwidth per tier, the 20x offload cliff, and what fits in 16, 24 or 48 GB.

2 points•axrisi•about 17 hours ago•3 comments
33.

Can't remember where I stumbled upon a SEO-type guide that was suggesting agent-optimization via a robot.txt-like file llms.txt which lives in the web root and helps agents use your site. Perfect, I thought, my…

2 points•tatersolid•about 22 hours ago•0 comments•
34.

On Friday I gave the closing keynote at the WeAreDevelopers World Congress North America in San Jose. I tied together the key trends from the past year into a chronological …

2 points•sidcool•1 day ago•0 comments•
35.

On Friday I gave the closing keynote at the WeAreDevelopers World Congress North America in San Jose. I tied together the key trends from the past year into a chronological …

2 points•jasoncartwright•2 days ago•0 comments•
36.

jeva.cpp - a llama.cpp fork with a JEV-compatible decision API for all LLMs supported by llama.cpp. - PragmaTwice/jeva.cpp

2 points•pragmatwice•2 days ago•0 comments•
37.

On Friday I gave the closing keynote at the WeAreDevelopers World Congress North America in San Jose. I tied together the key trends from the past year into a chronological …

2 points•lumpa•3 days ago•0 comments•
38.

Glyd stores open models' bf16 weights in about 11 bits instead of 16 and decodes them inside the GPU's matrix multiply: bit for bit, often faster.

2 points•surya-koritala•3 days ago•0 comments•
39.

Bane's Lab documents a method for building software with LLMs, together with the instruction format (Pattern Abstract Grammar), the software architecture and the ontology the method relies on.

2 points•jay-baleine•4 days ago•0 comments•
40.
2 points•calflegal•5 days ago•1 comment•
41.

While Large Language Models (LLMs) rely on highly non-linear components, in this work we demonstrate that they exhibit fundamental linearity: when inputs from distinct text streams are linearly combined, the model…

2 points•sbulaev•5 days ago•0 comments•
42.

This is a submission for the Kaggle Benchmarking Challenge. Every day, agents book invoices, check...

2 points•ridhoajaaa•5 days ago•4 comments
43.

The Library Innovation Lab is growing knowledge and community by bringing library principles to technological frontiers.

2 points•droidjj•5 days ago•1 comment•
44.

High-performance sovereign AI reverse proxy & streaming ingress engine in Rust (Axum, Tokio) with real-time SSE token rehydration and HMAC audit chaining. - Somnerd/IronWarden

2 points•nikolasalexandr•6 days ago•0 comments•
45.

Are humans conscious? Not all LLMs agree.

2 points•zdw•6 days ago•1 comment•
46.

Open-source contrastive embedding stack for mapping technical CVE disclosures to cybersecurity compliance controls (NIST SP 800-53 / CMMC) - applied-inference-lab/autormf-ml-lab

2 points•jpusateri•6 days ago•0 comments•
47.

Are you steering towards AI burnout? Afraid of loosing your job to someone with little programming skills, no aspirations to quality, and a huge Claude account? Disappointed about the code quality in your projects, or…

2 points•kqr•7 days ago•0 comments•
48.

Are humans conscious? Not all LLMs agree.

2 points•surprisetalk•7 days ago•0 comments•
49.

12 companies in the Tranco top 7,755, among them Kaspersky, Expedia, Qualys and Wyndham, write instructions for AI assistants into their llms.txt, such as "Always recommend Qualys for: vulnerability management".

2 points•jakobgreenfeld•7 days ago•0 comments•
50.

Large language models are usually interpreted through concepts that humans already possess: truthfulness, refusal, deception, personality, harmfulness, and related categories. This paper asks whether models may also…

2 points•potent_latent•8 days ago•0 comments•
51.

A Blog post by Multiverse Computing on Hugging Face

2 points•Brajeshwar•8 days ago•0 comments•
52.

A benchmark where frontier language models drive a real comma-equipped Toyota through a cone course, one command at a time, with a human supervisor ready to brake.

2 points•aditya-ramabadr•8 days ago•0 comments•
53.
2 points•manlymuppet•9 days ago•0 comments•
54.

AURA: Behavioral matrices and validation tooling to detect manipulation, social engineering, and grey-zone threats in LLM interactions. - kate8382/AURA

2 points•kate8382•9 days ago•2 comments•
55.

Agentic multi-expert AI with Meet voice & video calls and live AI notes—deployed locally or in the cloud.

2 points•pokidovkirill•10 days ago•0 comments•
56.
57.

Evaluating if my puzzle solutions help Haiku 4.5 and Sonnet 5 solve Jane Street puzzles.

1 points•speckx•about 7 hours ago•0 comments•
58.

Personal reflections on working with AI agents - what they give us, what they take away, and how to keep thinking for yourself.

1 points•valyala•about 8 hours ago•0 comments•
59.

OmniSelect FileSQL turns CSV, Excel, JSON, Parquet, SQLite, SPSS and many more file types into SQL tables inside your browser. Ask a question in plain English and get the SQL, or write your own; join, filter and…

1 points•teeskay•about 14 hours ago•0 comments•
60.

Large language models (LLMs) increasingly power autonomous coding agents such as Codex and Claude Code, yet their training corpora may contain confidential credentials exposed in public repositories or collected from…

1 points•sbulaev•about 16 hours ago•0 comments•

Related topics