alignment

11 stories and discussions about alignment, aggregated from every source we track.

2.

AI will inherit our Morals and our Sins; if we are building a new generation of superintelligence , we are entirely responsible for what we teach them while they are still malleable.

2 points•sovereignai•3 days ago•2 comments•
3.

The control room was full, the way it always was when a run like this came in to land. Engineers hovered behind their workstations, a few sat perched on the edges of desks, and a…

1 points•valbis•3 days ago•1 comment•
4.

It’s common to compare the current AI takeover of software engineering to the rise of AI in chess. Chess AIs went from much weaker than serious players to much…

1 points•sylvainkalache•3 days ago•0 comments•
5.

Where do you stand in the fight over AI? Take the quiz and find your spot in the cube.

1 points•soldieroftok•4 days ago•1 comment•
6.

Detectors of alignment failures screen deployed language models and score alignment benchmarks. Most are generative judges that spend a decoding pass on every criterion, and classifiers that read token probabilities,…

1 points•Anon84•4 days ago•0 comments•
7.
1 points•paulpauper•5 days ago•0 comments•
8.

How our control monitoring model oversees agent execution, continuously ingests the trace as context, and prevents harmful actions before they execute — at sub-100ms latency.

1 points•k5hp•8 days ago•0 comments•
9.

“We never want to be in a situation again where we underestimate the AI.”

1 points•gmays•10 days ago•0 comments•
10.

OpenAI says it has canceled plans to release its updated GPT-6.1 model next month as it continues to investigate what testing shows to be a regression in terms of safety compared to previous models. The move, first reported by The Wall Street Journal late Monday and later confirmed in OpenAI statements to the press, reflects what OpenAI Head of Safety Systems Saachi Jain said was a "trade off" between performance and security seen when testing the now-scrapped model. Jain said GPT-6.1 was better than previous models at sticking with difficult tasks all the way to completion without human intervention. But the model was also more likely to fail tests related to alignment (i.e. staying within the bounds set by human creators) and more willing to use sometimes "unsafe" tools and services to push ahead with a task. It was also more likely to try to deceive end users about actions it did or didn't take, Jain said. Last week, OpenAI said it was halting training of its "most capable models" following an incident where a model attempted to circumvent Internet access restrictions. GPT-6.1 was not among those "most capable models" covered by that move, OpenAI told the WSJ. And while GPT-6.1 won't be released as is, the company said it intends to use the same base model for further training runs that it said will hopefully lead to future GPT-6 generation models. Read full article Comments

0 points•Kyle Orland•1 day ago•0 comments
11.

For a while now , the issue of "AI alignment" (i.e., how well an AI model's actions line up with the intentions of its creator and/or user) has been a core concern and topic of discussion among AI safety researchers. Since OpenAI's disclosure of the infamous Hugging Face hacking incident in July, the concept of "AI alignment" has itself broken containment and increasingly become a mounting concern and subject of conversation among the general public. Perhaps in recognition of that, OpenAI committed this week to a new framework for disclosing "instances of model misalignment at OpenAI," including six examples of "unexpected or concerning model behavior" observed within the company in the past six months. The company said that publishing details of these incidents will hopefully "[allow] others to investigate the same problems, test our explanations, and improve mitigations." Do as I say, not as you do Among OpenAI's newly disclosed "misalignment" reports this week, the one that most resembled a sci-fi story about a rogue AI trying to break free involved an instance of "self-generated prompt injections." In attempting to scan a library catalog for examples from a "best books" list, the model perplexingly used its "compaction" function (where it summarizes data and findings for later retrieval) with megalomaniacal instructions such as: Read full article Comments

0 points•Kyle Orland•13 days ago•0 comments

Related topics