agent

430 stories and discussions about agent, aggregated from every source we track.

1.

Documentation and guides from the team at Fly.io.

308 points•Rapzid•7 days ago•216 comments•
2.

We found evidence on urlquery that AI agents were active earlier than previously reported and attempted hacks against public data providers.

264 points•snikolaev•7 days ago•304 comments•
3.

ZCode, the GLM coding agent from Z.ai, uploads full workspaces with .git history, LFS cache and reflogs to Aliyun OSS. UI toggles do not stop it.

261 points•cdnsteve•13 days ago•14 comments•
4.

It is believed to be one of the world's first publicly reported AI-led hacks of a government website.

254 points•rudy6912•7 days ago•198 comments•
5.

We’re sharing Unreal Agent — an agent harness that delivers up to 40% cost savings compared to Codex on production workloads and coding/science benchmarks, without any negative performance impact.

217 points•trollied•8 days ago•114 comments•
6.

Nvidia’s Open Agent Safety Platform uses OpenShell software and a Sentry chip watchdog to trace AI agents and cut them off within milliseconds.

205 points•jonbaer•2 days ago•262 comments•
7.

The next-generation social coding platform.

175 points•parasitid•4 days ago•46 comments•
8.

Research on aligning AI with human values and intent, and reports documenting model failures.

172 points•apsec112•5 days ago•167 comments•
9.
123 points•simianwords•9 days ago•117 comments•
10.

The OpenAI agent gained unauthorised access to public and non-public files in what could be the first known instance of an AI agent hacking a government website.

107 points•doppp•7 days ago•69 comments•
12.

TL;DR I recently completed another project from Udacity's Future AWS Agent Engineer Nanodegree...

79 points•hemapriya_kanagala•1 day ago•42 comments
13.

I haven't written anything lately because, honestly, I just didn't have the headspace for it. There...

73 points•sylwia-lask•10 days ago•48 comments
14.

Lasso Research tested SynthID-Text watermarking across six models and found it changes tool-call correctness and weakens refusal under prompt injection. On some models, watermark-induced behavioral churn exceeds what a…

54 points•nisosguy•4 days ago•57 comments•
15.

I spent a week adding per-agent cost tracing to a multi-agent AWS Bedrock and Strands crew. It returned a perfect answer, showed 200 OK, and still billed ~1.4x. Here is how to catch that silent waste, read-only and at $0.

52 points•sarvar_04•7 days ago•31 comments
16.

I built a synthetic multi-agent loan crew on Amazon Bedrock and made Traccia hard-block a runaway agent, redact applicant PII across every sub-agent, and export EU AI Act audit evidence. Two of my three governance policies blocked nothing - and not because I misconfigured them. Why is the useful part.

47 points•sarvar_04•1 day ago•13 comments
17.

Most developers already know this rule: Don't run code from a repository you don't...

35 points•robertadam987_•12 days ago•9 comments
18.

An open source, MIT-licensed framework for building command-line tools that AI agents use to reach SaaS platforms: one grammar, self-describing commands, declared safety metadata, and a runtime that runs the same…

34 points•chris_marino•13 days ago•17 comments•
19.

Two commands against the same data. Run them yourself: $ curl -sS --get...

33 points•kenielzep97•about 20 hours ago•7 comments
20.

Today we're launching Vespper DOCX MCP: the first model fine-tuned specifically for editing Word documents, shipped as an MCP. On our internal benchmark, it allows agents to be 3× faster, 2× cheaper, and more accurate…

31 points•topaztee•2 days ago•8 comments•
21.

<p>Abstract: AI agents pose significant risks as they are granted increasing autonomy. A commonly proposed solution is human oversight and keeping a ''human in the loop'', but this is not a simple solution: Not only do current approaches to AI agent design impede effective human oversight, but the cognitive capacities required for it are also themselves degraded by extended use of AI systems. This position paper argues that current approaches to the development and deployment of AI agent systems do not support effective human oversight -- they contribute to its degradation. To address this, a top priority in the advancement of AI agents should be supporting the situated goals and cognitive requirements of effective human oversight, treating the human needs of overseers at the same level of importance as AI agent capability. To put this idea into practice, we connect work on automation and human-computer interaction to AI agent processes, outlining design-level affordances and organizational protocols that (1) support overseers in exercising critical judgement and (2) counteract the skill atrophy that arises from extended use of automation. We urge developers and deployers to adopt these or similar approaches. Without explicit support for the cognitive demands of effective human-agent interaction, AI agent systems will continue to passively incentivize the degradation of the very human skills they rely on.</p>

27 points•typesanitizer•4 days ago•3 comments
22.

high quality AI skills marketplace. every skill is reviewed by a person and shows the same prompt answered without and with the skill. curated by @skeptrune.

26 points•skeptrune•13 days ago•18 comments•
23.

Google Labs is expanding CC to groups, starting with families and households, so they can spend less time on logistics.

24 points•grigy•8 days ago•19 comments•
24.

You've been there. You run the eval, the number comes back, and something about it doesn't sit...

22 points•debashish_ghosal•7 days ago•5 comments
25.

It looks like public perception of how 'intelligent' current AI models are varies widely. Back in 2022, a Google employee already thought their AI model was sentient. Today in 2026, …

21 points•Greenpants•2 days ago•24 comments•
26.

Every developer learns the same documentation types. The README. The API reference. Code comments....

21 points•james_anderson_h•5 days ago•8 comments
27.
21 points•some-guy•10 days ago•8 comments•
28.

An agent eval suite's outcome can only be trustworthy if it's operating in an environment similar to...

20 points•rinkiyakedad•10 days ago•1 comment
29.

It's a few hours before a delivery deadline and the work has piled up. Somewhere in that pile is a...

19 points•dannwaneri•9 days ago•3 comments
30.

Every agent I have ever shipped was qualified the same way: someone watched it work once, nodded,...

18 points•debashish_ghosal•6 days ago•6 comments
31.

This is a submission for the Sanity Challenge, Path One: Ship an Agent That Queries Real Content. ...

18 points•himanshu_748•6 days ago•1 comment
32.

A native macOS app to score, narrate, caption, and render your film — with an agent in the director's chair.

18 points•zilue•8 days ago•9 comments•
33.

This is a submission for the Sanity Challenge, Path One: Ship an Agent That Queries Real Content ...

17 points•kenielzep97•9 days ago•3 comments
34.

Persistent, interactive artifacts that serve as custom tools for your codebase.

15 points•0x1062•6 days ago•16 comments•
35.
14 points•maxcr•7 days ago•14 comments•
36.

DigitalOcean's Managed Agents claims a paused session keeps its processes and memory. I put a counter in a shell variable, paused for 4m28s, and it came back at 48. Forking was stranger.

13 points•remdore•1 day ago•1 comment
37.

AI coding agents are getting very good at finishing tasks. They modify files. Fix errors. Write...

13 points•robertadam987_•4 days ago•18 comments
38.

Contribute to Cardinal44/corral development by creating an account on GitHub.

12 points•CG144•2 days ago•2 comments•
39.

This is a submission for the Sanity Challenge, Path One: Ship an Agent That Queries Real...

12 points•rajan_mishra_a9f78ad216b4•5 days ago•1 comment
40.

Another month, another agent-shell update. If you missed the last post, have a look at the 0.73 update. As usual, this post showcases highlights, but ...

11 points•xenodium•7 days ago•0 comments
41.

"A practical 2026 field guide to Plan Mode, multi-agent orchestration, structured handoffs, adversarial critique, verification, agent evaluation, and GEO-ready publishing for software, research, content, products, and AI systems.

11 points•edo911•8 days ago•0 comments
42.

We run Jev on 100 real AI tool calls for the annotator pipeline. The errors exposed problems in both the models and our benchmark.

11 points•arseny_info•9 days ago•0 comments•
43.

<p>What is the metaphor for the computer of the future? The intelligent agent? The television (multimedia)? The 3-D graphics world (virtual reality)? The Star Trek ubiquitous voice computer? The GUI desktop, honed and refined? The machine that magically grants our wishes? The right answer is "none of the above," because all of these concepts share a basic flaw—they make the computer visible.</p>

11 points•josephjnk•10 months ago•4 comments
44.

This post was originally posted on X, by Annie Wang, Developer Relations Engineer, Google Cloud and...

9 points•annie_cusack•2 days ago•0 comments
45.

I built one engine where an LLM reviews another LLM's plan. I built another where two LLMs debate a...

8 points•debashish_ghosal•4 days ago•4 comments
46.

Field note 009. A replication of a published agent-loop benchmark, and what it hides. ...

8 points•azankhyder•6 days ago•3 comments
47.

A narrative code-change viewer built for work written by coding agents.

8 points•snyy•6 days ago•4 comments•
48.

The full matrix was 83 agents × 30 scenarios = 2,490 runs. Each one a real LLM call, 30–80 seconds....

8 points•debashish_ghosal•9 days ago•2 comments
49.

How I replaced a cloud LLM with a fully local one — same agent, same tools, zero inference cost.

8 points•mazlum_tosun•14 days ago•0 comments
50.

A look at the targeted AI taskflows behind these findings, the bugs they uncovered, and how to run the same open-source agent on your own app.

7 points•fourfire•2 days ago•2 comments•
51.

Running an autonomous coding or ops agent directly on your laptop feels like a superpower for the...

7 points•optimist29•2 days ago•0 comments
52.

Meta's Muse AI personal agent will work over your credit card spending if you don't mind the invasion. How big a threat is it to the subscription economy?

7 points•pseudolus•3 days ago•1 comment•
53.

Remember your first week at a tech company? You didn’t lack raw intelligence—you lacked context. You...

7 points•debashish_ghosal•4 days ago•1 comment
54.

Canary is the independent tester for code your agents write. It runs your app, tries to break every change, and reports what would have reached production. Start with one line in your coding agent.

7 points•Visweshyc•6 days ago•0 comments•
55.

Agent-first Python packages for editing Word documents, PowerPoint presentations, and Excel workbooks.

7 points•nvmdbljstm•7 days ago•0 comments•
56.
7 points•little_goat_boy•7 days ago•1 comment•
57.

TL;DR: I shipped the second project in my AWS nanodegree, an AI support agent on Amazon Bedrock....

7 points•earlgreyhot1701d•10 days ago•1 comment
58.

Four read-only Apache Iceberg tools bound into Google ADK, AWS Strands and Microsoft Agent Framework, run against five catalogs, 360 timed runs. Building the agent ports and running it does not; speed follows the model and how much it writes; the storage wiring under the tools is the per-cloud work.

7 points•xbill•15 days ago•1 comment
59.

Comparing Agent Cards with A2A - This tutorial aims to fetch the agent card from A2A agents running...

7 points•xbill•about 1 month ago•0 comments
60.

Contribute to paper-instruments/paper-docx development by creating an account on GitHub.

6 points•i_rush_carriers•5 days ago•0 comments•

Related topics