agent
430 stories and discussions about agent, aggregated from every source we track.
Documentation and guides from the team at Fly.io.
We found evidence on urlquery that AI agents were active earlier than previously reported and attempted hacks against public data providers.
ZCode, the GLM coding agent from Z.ai, uploads full workspaces with .git history, LFS cache and reflogs to Aliyun OSS. UI toggles do not stop it.
It is believed to be one of the world's first publicly reported AI-led hacks of a government website.
We’re sharing Unreal Agent — an agent harness that delivers up to 40% cost savings compared to Codex on production workloads and coding/science benchmarks, without any negative performance impact.
Nvidia’s Open Agent Safety Platform uses OpenShell software and a Sentry chip watchdog to trace AI agents and cut them off within milliseconds.
The next-generation social coding platform.
Research on aligning AI with human values and intent, and reports documenting model failures.
The OpenAI agent gained unauthorised access to public and non-public files in what could be the first known instance of an AI agent hacking a government website.
TL;DR I recently completed another project from Udacity's Future AWS Agent Engineer Nanodegree...
I haven't written anything lately because, honestly, I just didn't have the headspace for it. There...
Lasso Research tested SynthID-Text watermarking across six models and found it changes tool-call correctness and weakens refusal under prompt injection. On some models, watermark-induced behavioral churn exceeds what a…
I spent a week adding per-agent cost tracing to a multi-agent AWS Bedrock and Strands crew. It returned a perfect answer, showed 200 OK, and still billed ~1.4x. Here is how to catch that silent waste, read-only and at $0.
I built a synthetic multi-agent loan crew on Amazon Bedrock and made Traccia hard-block a runaway agent, redact applicant PII across every sub-agent, and export EU AI Act audit evidence. Two of my three governance policies blocked nothing - and not because I misconfigured them. Why is the useful part.
Most developers already know this rule: Don't run code from a repository you don't...
An open source, MIT-licensed framework for building command-line tools that AI agents use to reach SaaS platforms: one grammar, self-describing commands, declared safety metadata, and a runtime that runs the same…
Two commands against the same data. Run them yourself: $ curl -sS --get...
Today we're launching Vespper DOCX MCP: the first model fine-tuned specifically for editing Word documents, shipped as an MCP. On our internal benchmark, it allows agents to be 3× faster, 2× cheaper, and more accurate…
<p>Abstract: AI agents pose significant risks as they are granted increasing autonomy. A commonly proposed solution is human oversight and keeping a ''human in the loop'', but this is not a simple solution: Not only do current approaches to AI agent design impede effective human oversight, but the cognitive capacities required for it are also themselves degraded by extended use of AI systems. This position paper argues that current approaches to the development and deployment of AI agent systems do not support effective human oversight -- they contribute to its degradation. To address this, a top priority in the advancement of AI agents should be supporting the situated goals and cognitive requirements of effective human oversight, treating the human needs of overseers at the same level of importance as AI agent capability. To put this idea into practice, we connect work on automation and human-computer interaction to AI agent processes, outlining design-level affordances and organizational protocols that (1) support overseers in exercising critical judgement and (2) counteract the skill atrophy that arises from extended use of automation. We urge developers and deployers to adopt these or similar approaches. Without explicit support for the cognitive demands of effective human-agent interaction, AI agent systems will continue to passively incentivize the degradation of the very human skills they rely on.</p>
high quality AI skills marketplace. every skill is reviewed by a person and shows the same prompt answered without and with the skill. curated by @skeptrune.
Google Labs is expanding CC to groups, starting with families and households, so they can spend less time on logistics.
You've been there. You run the eval, the number comes back, and something about it doesn't sit...
It looks like public perception of how 'intelligent' current AI models are varies widely. Back in 2022, a Google employee already thought their AI model was sentient. Today in 2026, …
Every developer learns the same documentation types. The README. The API reference. Code comments....
An agent eval suite's outcome can only be trustworthy if it's operating in an environment similar to...
It's a few hours before a delivery deadline and the work has piled up. Somewhere in that pile is a...
Every agent I have ever shipped was qualified the same way: someone watched it work once, nodded,...
This is a submission for the Sanity Challenge, Path One: Ship an Agent That Queries Real Content. ...
A native macOS app to score, narrate, caption, and render your film — with an agent in the director's chair.
This is a submission for the Sanity Challenge, Path One: Ship an Agent That Queries Real Content ...
Persistent, interactive artifacts that serve as custom tools for your codebase.
DigitalOcean's Managed Agents claims a paused session keeps its processes and memory. I put a counter in a shell variable, paused for 4m28s, and it came back at 48. Forking was stranger.
AI coding agents are getting very good at finishing tasks. They modify files. Fix errors. Write...
Contribute to Cardinal44/corral development by creating an account on GitHub.
This is a submission for the Sanity Challenge, Path One: Ship an Agent That Queries Real...
Another month, another agent-shell update. If you missed the last post, have a look at the 0.73 update. As usual, this post showcases highlights, but ...
"A practical 2026 field guide to Plan Mode, multi-agent orchestration, structured handoffs, adversarial critique, verification, agent evaluation, and GEO-ready publishing for software, research, content, products, and AI systems.
We run Jev on 100 real AI tool calls for the annotator pipeline. The errors exposed problems in both the models and our benchmark.
<p>What is the metaphor for the computer of the future? The intelligent agent? The television (multimedia)? The 3-D graphics world (virtual reality)? The Star Trek ubiquitous voice computer? The GUI desktop, honed and refined? The machine that magically grants our wishes? The right answer is "none of the above," because all of these concepts share a basic flaw—they make the computer visible.</p>
This post was originally posted on X, by Annie Wang, Developer Relations Engineer, Google Cloud and...
I built one engine where an LLM reviews another LLM's plan. I built another where two LLMs debate a...
Field note 009. A replication of a published agent-loop benchmark, and what it hides. ...
A narrative code-change viewer built for work written by coding agents.
The full matrix was 83 agents × 30 scenarios = 2,490 runs. Each one a real LLM call, 30–80 seconds....
How I replaced a cloud LLM with a fully local one — same agent, same tools, zero inference cost.
A look at the targeted AI taskflows behind these findings, the bugs they uncovered, and how to run the same open-source agent on your own app.
Running an autonomous coding or ops agent directly on your laptop feels like a superpower for the...
Meta's Muse AI personal agent will work over your credit card spending if you don't mind the invasion. How big a threat is it to the subscription economy?
Remember your first week at a tech company? You didn’t lack raw intelligence—you lacked context. You...
Canary is the independent tester for code your agents write. It runs your app, tries to break every change, and reports what would have reached production. Start with one line in your coding agent.
Agent-first Python packages for editing Word documents, PowerPoint presentations, and Excel workbooks.
TL;DR: I shipped the second project in my AWS nanodegree, an AI support agent on Amazon Bedrock....
Four read-only Apache Iceberg tools bound into Google ADK, AWS Strands and Microsoft Agent Framework, run against five catalogs, 360 timed runs. Building the agent ports and running it does not; speed follows the model and how much it writes; the storage wiring under the tools is the per-cloud work.
Comparing Agent Cards with A2A - This tutorial aims to fetch the agent card from A2A agents running...
Contribute to paper-instruments/paper-docx development by creating an account on GitHub.