llms
99 stories and discussions about llms, aggregated from every source we track.
Are you steering towards AI burnout? Afraid of loosing your job to someone with little programming skills, no aspirations to quality, and a huge Claude account? Disappointed about the code quality in your projects, or…
Run Q4 small language models entirely in the browser with WebGPU. Chat, benchmark, and compare PetitGPT, SmolLM2, and friends on-device.
AI drastically shortened its design time; it will only get faster
I was intrigued by Jev and the self-hostable projects appearing around it, such as OpenJev and SemIf . Reading about them introduced me ...
Are you steering towards AI burnout? Afraid of loosing your job to someone with little programming skills, no aspirations to quality, and a huge Claude account? Disappointed about the code quality in your projects, or…
You can't get them to stop talking about spines!
What the independent measurements of TypeSafe's Jev found in its first eight days: arXiv preprints, GitHub evaluations and blog benchmarks, each traced to its primary source. Accuracy, calibration, speed, cost, failure modes, the prior art, the open alternatives, and what is still unmeasured.
An esoteric programming language built to break LLMs and coding agents.
This is a submission for the Kaggle Benchmarking Challenge What I Benchmarked ...
The framework for programming—rather than prompting—language models.
Every Apple Silicon Mac can run AI locally. Here's what model sizes actually fit your machine, the tools to use, and the daily-driver use nobody talks about.
A from-scratch derivation of REINFORCE for language models
Plus Claude Opus 5.5, GPT-6 Sol, GPT-6 Luna, and a new price war
StarSkirmish: where LLMs write code to play StarCraft.
"Millions of people around the world will soon consider large models ‘hoovering up’ all their work to be an astonishing theft of unprecedented proportions."
Using "think" to describe what an LLM does reminds me of the 16th century, when astronomers didn't really believe in the heliocentric model, but used it for calculations anyway because the math was simpler. We'll start…
Benchmark of TypeSafe's Jev against Sonnet 5, GPT-5 nano and local LLMs on 770 Reddit AITA verdicts: Brier scores, latency and cost - dchristopoulos/jev-aita
The model deciding whether to wake someone up doesn't need to sing Bohemian Rhapsody in Klingon. Working with Jev has me thinking about how much less we should ask of AI.
Some miscellaneous thoughts on what LLMs are changing in my field.
AI drastically shortened its design time; it will only get faster
VRAM for LLMs is a bandwidth problem: every token streams the whole model from memory. Bandwidth per tier, the 20x offload cliff, and what fits in 16, 24 or 48 GB.
Can't remember where I stumbled upon a SEO-type guide that was suggesting agent-optimization via a robot.txt-like file llms.txt which lives in the web root and helps agents use your site. Perfect, I thought, my…
On Friday I gave the closing keynote at the WeAreDevelopers World Congress North America in San Jose. I tied together the key trends from the past year into a chronological …
On Friday I gave the closing keynote at the WeAreDevelopers World Congress North America in San Jose. I tied together the key trends from the past year into a chronological …
jeva.cpp - a llama.cpp fork with a JEV-compatible decision API for all LLMs supported by llama.cpp. - PragmaTwice/jeva.cpp
On Friday I gave the closing keynote at the WeAreDevelopers World Congress North America in San Jose. I tied together the key trends from the past year into a chronological …
Glyd stores open models' bf16 weights in about 11 bits instead of 16 and decodes them inside the GPU's matrix multiply: bit for bit, often faster.
Bane's Lab documents a method for building software with LLMs, together with the instruction format (Pattern Abstract Grammar), the software architecture and the ontology the method relies on.
While Large Language Models (LLMs) rely on highly non-linear components, in this work we demonstrate that they exhibit fundamental linearity: when inputs from distinct text streams are linearly combined, the model…
This is a submission for the Kaggle Benchmarking Challenge. Every day, agents book invoices, check...
The Library Innovation Lab is growing knowledge and community by bringing library principles to technological frontiers.
High-performance sovereign AI reverse proxy & streaming ingress engine in Rust (Axum, Tokio) with real-time SSE token rehydration and HMAC audit chaining. - Somnerd/IronWarden
Are humans conscious? Not all LLMs agree.
Open-source contrastive embedding stack for mapping technical CVE disclosures to cybersecurity compliance controls (NIST SP 800-53 / CMMC) - applied-inference-lab/autormf-ml-lab
Are you steering towards AI burnout? Afraid of loosing your job to someone with little programming skills, no aspirations to quality, and a huge Claude account? Disappointed about the code quality in your projects, or…
Are humans conscious? Not all LLMs agree.
12 companies in the Tranco top 7,755, among them Kaspersky, Expedia, Qualys and Wyndham, write instructions for AI assistants into their llms.txt, such as "Always recommend Qualys for: vulnerability management".
Large language models are usually interpreted through concepts that humans already possess: truthfulness, refusal, deception, personality, harmfulness, and related categories. This paper asks whether models may also…
A Blog post by Multiverse Computing on Hugging Face
A benchmark where frontier language models drive a real comma-equipped Toyota through a cone course, one command at a time, with a human supervisor ready to brake.
AURA: Behavioral matrices and validation tooling to detect manipulation, social engineering, and grey-zone threats in LLM interactions. - kate8382/AURA
Agentic multi-expert AI with Meet voice & video calls and live AI notes—deployed locally or in the cloud.
Evaluating if my puzzle solutions help Haiku 4.5 and Sonnet 5 solve Jane Street puzzles.
Personal reflections on working with AI agents - what they give us, what they take away, and how to keep thinking for yourself.
OmniSelect FileSQL turns CSV, Excel, JSON, Parquet, SQLite, SPSS and many more file types into SQL tables inside your browser. Ask a question in plain English and get the SQL, or write your own; join, filter and…
Large language models (LLMs) increasingly power autonomous coding agents such as Codex and Claude Code, yet their training corpora may contain confidential credentials exposed in public repositories or collected from…