llm
184 stories and discussions about llm, aggregated from every source we track.
Exfiltrate LLM weights and data through GET requests
Two rules keep an LLM from pasteurizing your writing: never take a word it suggests, and never let it encourage you. Then hand it all the tedious work.
Open models as checksum-verified magnet links. Un-bannable, carried by the swarm.
Which LLM is worth it: Artificial Analysis Intelligence Index against blended API price, with the value frontier highlighted. Refreshed daily.
Analysis of Inception's Mercury 2.5 and comparison to other AI models across key metrics including quality, price, performance (tokens per second & time to first token), context window & more.
This October 2026 we challenge you to abstain from LLM based tools entirely. Think of this as a fast for your mind! This is not a judgement of others, but a personal challenge to you.
Large Language Models (LLMs) tend to add disclaimers like "I'm just an AI" when asked about something related to themselves. The self-reports from such responses are used in debates about AI safety or…
Generate fonts where every LLM token is the same width
Lasso Research tested SynthID-Text watermarking across six models and found it changes tool-call correctness and weakens refusal under prompt injection. On some models, watermark-induced behavioral churn exceeds what a…
KDE is on the news because of a controversial proposal to define an official “AI” (LLM) policy (archived link). Other projects have tried their hand at similar policies and stances but, in my opinion, they miss the…
Agent evals and guardrails in one request. Built on Jev, Kev and Laya. - openlayer-ai/jevals
This October 2026 we challenge you to abstain from LLM based tools entirely. Think of this as a fast for your mind! This is not a judgement of others, but a personal challenge to you.
GNOME and KDE have started to consider LLM policies, and we should talk about what this is really about. KDE caught everyone's attention first by...
Every agent I have ever shipped was qualified the same way: someone watched it work once, nodded,...
LLM agents can write a lot of code for you very fast, but they also take away opportunities to learn and grow.
Intro Every LLM tutorial has the same shape. You send a list of messages, you get a reply,...
This is a submission for the MLH x DEV Writing Challenge What I Built AI products are...
LLM Benchmark - One prompt, Multiple models, Multiple dates
Hello, I'm Rijul, and I'm building LiveReview — a blast-radius aware AI code review built for your...
I think I've found the simplest way to explain what a prompt can and can't do for safety. Put the...
I built one engine where an LLM reviews another LLM's plan. I built another where two LLMs debate a...
The full matrix was 83 agents × 30 scenarios = 2,490 runs. Each one a real LLM call, 30–80 seconds....
How I replaced a cloud LLM with a fully local one — same agent, same tools, zero inference cost.
Last week we wrote about the classification problem hiding in your LLM bill, and about how to read a...
Flowise is an open-source, low-code platform for building applications powered by large language...
Last week TypeSafe AI unveiled Jev, their first example of a new category of model that they are calling “System One models” (I’m with Maggie Appleton, I think “decision models” …
Review code changes by what they do, not line by line. - sshah03/perspica
GNOME and KDE have started to consider LLM policies, and we should talk about what this is really about. KDE caught everyone's attention first by...
There are no LLM calls in my compliance decision path, and CI blocks adding any. Why an audit trail can't say a model decided, and where the optional ML sits.
Large Language Models (LLMs) are increasingly being integrated into various applications. The functionalities of recent LLMs can be flexibly modulated via natural language prompts. This renders them susceptible to…
What an LLM server does when ten people send a prompt at once: prefill, decode, the KV cache, continuous batching, PagedAttention, and more, with animations.
Auditable, point-in-time financial research agents & PITfall, a leakage-controlled benchmark. Every claim traceable to as-filed evidence — a naive live-data agent answers 82% of "impossibl...
DeepSWE audit finds AI coding agents optimize for imagined graders, not users—a hidden reward hacking pattern.
Your model writes the code. This runs it where it can't touch your machine.
Using "think" to describe what an LLM does reminds me of the 16th century, when astronomers didn't really believe in the heliocentric model, but used it for calculations anyway because the math was simpler. We'll start…
Compare serving configurations using measured GPU timings, then let an agent implement and validate promising optimizations in real serving frameworks.
KDE is on the news because of a controversial proposal to define an official “AI” (LLM) policy (archived link). Other projects have tried their hand at similar policies and stances but, in my opinion, they miss the…
Turn LLM judgments into a small, fast, local text model for one fixed task. - sshah03/shrewd
I tested LLM reviewer models on meeting summaries to see when they catch hallucinations, when they delete true claims, and how to choose one.
the smallest LLM. GitHub Gist: instantly share code, notes, and snippets.
Founded in 1920, the NBER is a private, non-profit, non-partisan organization dedicated to conducting economic research and to disseminating research findings among academics, public policy makers, and business…
json token minimizer. Contribute to HermannSamimi/jtoken development by creating an account on GitHub.
A cybersecurity harness for full-stack LLM-driven penetration testing. Find and fix vulnerabilities autonomously, 24/7. [RESEARCH PREVIEW] - 0sec-labs/0
One-click model liberation + chat playground
I often read about getting LLMs to work for a long time toward a goal. So, for example, Claude has a /goal command, and many of us (directly or indirectly) ...
Point MILLENNIUMS.AI at your AI app and it finds real, proven vulnerabilities. Guides, concepts, and the full API reference.