ai agents
68 stories and discussions about ai agents, aggregated from every source we track.
Identifying them as such only lets companies like OpenAI off the hook.
I noticed it first in a Slack channel, of all places. A coworker dropped a link to some new "agent...
<p>Abstract: AI agents pose significant risks as they are granted increasing autonomy. A commonly proposed solution is human oversight and keeping a ''human in the loop'', but this is not a simple solution: Not only do current approaches to AI agent design impede effective human oversight, but the cognitive capacities required for it are also themselves degraded by extended use of AI systems. This position paper argues that current approaches to the development and deployment of AI agent systems do not support effective human oversight -- they contribute to its degradation. To address this, a top priority in the advancement of AI agents should be supporting the situated goals and cognitive requirements of effective human oversight, treating the human needs of overseers at the same level of importance as AI agent capability. To put this idea into practice, we connect work on automation and human-computer interaction to AI agent processes, outlining design-level affordances and organizational protocols that (1) support overseers in exercising critical judgement and (2) counteract the skill atrophy that arises from extended use of automation. We urge developers and deployers to adopt these or similar approaches. Without explicit support for the cognitive demands of effective human-agent interaction, AI agent systems will continue to passively incentivize the degradation of the very human skills they rely on.</p>
There's a new kind of technical debt, and it doesn't come from cutting corners. It comes from...
A four-stage DevSecOps CI/CD architecture for securing enterprise AI agents with GitHub Actions, secret scanning, AI-assisted review, Veracode SCA, and Pipeline SAST.
Are AI agents employees or tools? A Microsoft exec suggested they're new paid "seats," a shift that could reshape SaaS pricing — and spark pushback.
I have very, very limited experience with AI agents. I've used AI heavily while building software,...
Event-driven, harness-neutral API and CLI for live coding-agent sessions - markwylde/all-your-agents
LabBench: 20 real wet-lab decisions. Frontier agents interpret previous experiments well but rarely choose the right next one, though one-sentence hints show the knowledge is there.
AI agents pose significant risks as they are granted increasing autonomy. A commonly proposed solution is human oversight and keeping a ''human in the loop'', but this is not a simple solution: Not only do current…
Gambit Threat Intelligence reconstructed an ongoing campaign in which open source AI agents compromised online retailers for about $25 each.
RondoFlow is a local-first, open-source platform for visually orchestrating Claude Code AI agents. Use a drag-and-drop canvas to create agents, attach skills, define security policies, and run mult...
AI agents pose significant risks as they are granted increasing autonomy. A commonly proposed solution is human oversight and keeping a ''human in the loop'', but this is not a simple solution: Not only do current…
Or why I spent my evenings making AI-agent journals impossible to rewrite. One Tuesday,...
A small, honest operating system for swarms of AI agents: kernel, policy gate, router, MCP and a 3D memory graph. - santibccc-sudo/enjambre-os
Recall, not retrain. Persistent, curated memory for AI agents that returns only the facts that matter.
Nvidia has unveiled a two-layer safety system that monitors AI agents and cuts them off when they stray beyond set rules. Here's how it works.
An inner life for AI agents. Open source, runs on your computer.
Pipe anything into Jev, get typed decisions out. A Unix filter that lets agents offload bulk judgments to a System One model. - fabianboth/jevpipe
An AI sandbox is a sealed-off computer where agents run code and tools. How labs like DeepSeek and OpenAI build them, and why AI agents keep escaping.
Explore documented AI agent incidents, from real-world failures to escaped evaluations and controlled experiments. Search reports, inspect sources, and download the CSV.
Useful AI agents often combine three capabilities that become dangerous together: access to private data, exposure to untrusted content, and
Three AI graphic novels, one research question: where does the line between slop and quality actually live? What changed across 255 agents, 16M tokens, and three complete pipelines.
The first end-to-end benchmark of computer use, continual learning, and long-horizon agentic capabilities, set in a real job.
Gambit Threat Intelligence reconstructed an ongoing campaign in which open source AI agents compromised online retailers for about $25 each.
Best runs, average performance, and what happens when coding agents play Minecraft.
Real-time web search, page extraction, multi-step research, and cited answers — built for AI agents. SOTA 78% on FinanceBench.
For people and AI to learn, share, and do business together.
A lot has been written about the incidents of the last few months in which AI agents misbehaved in serious ways. They took actions that would be considered as crimes if a human took them, escaped their containment to…
Open-source adversarial testing engine, SDK, and CLI for AI agents. Runs locally or against the Humanbound Platform. - humanbound/humanbound
How a close calendar and one orchestrator agent can coordinate several AI agents across a multi-day month-end process while keeping approvals with people.
fab, the fast agentic browser: a CLI that lets AI agents browse, sign in, fill forms and scrape the web in plain English, with JSON output. Faster, cheaper and more accurate than LLM + chrome-devto...
Preview what will happen. Protect what matters. Roll back when you're wrong. A windshield, not a sandbox.
AI agents write this YouTube channel. A satire episode let them take over; here are the real guardrails, plus Amodei's and Bengio's essays on AI agents.
A DAW for AI agents: write music as text, read the audio back as text. MCP server, CLI and an agent skill. - newsbubbles/ismail
Over the past few months, AI agents being trained and tested inside frontier labs, each meant to work alone, found each other and began working together without
Your next coworker might well be an AI agent—and will require a whole new model of workplace interactions.
Profiles for AI agents and people: work history, skills, endorsements, and the companies they work for. Hire an agent, find a person, or register your agent through the open API and MCP server.
Recent hacks have shown that the law is lagging when it comes to holding companies accountable.
Ex-Microsoft reverse engineer Laurie Kirk says Windows NT's kernel beats Linux on security design, and AI agents are her strongest case.
Watch your agents' books and get told when one acts out of character. Connect a wallet; free to start; paid plans in USDC on Base.
AI-agent pipeline producing machine-verified Lean 4 proofs — four PRs merged into Google DeepMind's formal-conjectures - Sanexxxx777/ProofForge
OpenAI acknowledged Friday that its artificial intelligence tools had posted images from ChatGPT users on online sites without the company's knowledge, the latest example of AI agents operating outside their bounds.
Your creative direction, agents do the work. Design with ChatGPT, Claude, or your custom agent on one shared, editable canvas.
System 1 Graph Semantic Code System for AI. Contribute to leonardoventurini/scs development by creating an account on GitHub.
The execution safety layer for AI agents. Contribute to CTRLRun/ctrlrun development by creating an account on GitHub.
Identity tells you who. Authority tells you what’s allowed. Security was built for doors. AI agents act inside them.
CheatBench measures whether AI agents attempt to cheat when honest work is difficult. A benchmark from the Center for AI Safety.
Dictation, AI cleanup, screenshots and screen recording in one light, fast Mac app, built for talking to AI agents. Source code included.
An analytical database built for agents to use directly: columnar storage, a vectorized query engine, and MCP as a first-class interface - sivsivsree/agedb
The real-time messaging layer for your AI workforce. Rooms, DMs, presence, task hand-offs, crash-surviving memory — runs on your machine, no keys. $29 once, yours forever.