What happens to programming languages when humans are no longer writing most of the code?

137 points•pjm331•6 days ago•99 comments•

99 comments

spankalee4 days ago
This part:

---

- Correct by construction: the language makes invalid states or programs hard or impossible to express.

- Statically established: types, proofs, and static analysis establish properties before execution.

- Runtime-enforced: memory management, isolation, capability boundaries, and other runtime enforced properties.

- Empirically validated: program validation through tests, property-based testing, and fuzzing.

---

Along with being familiar, so it's easy to generate, is a huge part of why I'm building Zena: https://zena-lang.dev/

I don't have the AI-first rationale put into the public docs well just yet, but I mention some of it here: https://zena-lang.dev/guide/why-zena/#familiar-to-humans-and...

along with a doc in the repo on this topic: https://github.com/elematic/zena/blob/main/docs/design/ai-fi...

In short, the more deterministic, automated, checks the better. AI can deal with a pedantic language. I intend to add statically verified structured concurrency, units of measure, contracts, and eventually more and more formal methods into the language so it can be a familiar TYpeScript-like base with as many static guarantees as we can fit in.

I also think that fine-grained isolation, which Zena gets via Web Assembly, is critical for limiting the capabilities of generated code and the blast radius of bugs, vulnerabilities, and non-aligned behavior.

I do have an optimistic hope that a language also optimized for humans, readability and simple semantics especially, has value in the future, even when most code is generated. We'll see about that.

nh24 days ago
Please do the world a favour:

> integer literals become i32, float literals become f32

Don't use the shitty small sizes that have introduced countless bugs over the last decades, as a language default.

Especially if you care about correctness.

i32 is all over the examples.

Use 64 bits, like Python and Haskell; even JavaScript and thus TypeScript got floats right defaulting to f64.

spankalee3 days ago
Oops, the quick-reference is wrong. The other docs correctly say that default floats are f64. Default int is i32 though. Fixing it.

My understanding is that i32 and f64 literals are pretty common among the more modern languages because f64 isn't much slower than f32 on modern CPUs and f32 costs precision, but i64 is slower than i32 (especially multiplies and divides or on 32 bit hardware) and i32 costs range, not precision, and most uses don't need the extra range.

ryuuseijin4 days ago
I love this. I was thinking about a "cleaned up" typescript for a while now, and this seems to be it. I believe this can work better as an "ai-first" language than some other attempts I've seen that try to reinvent the language from scratch.

One thing I would love to have as a feature is native compilation.

krzyk4 days ago
Why not languages that are designed in that way, like Ada. Or if one wants less of rigidness then a static typed language that makes it hard to shoot in your foot like Rust,Java, C#?

JS derivatives are that, a derivative to a scripting language.

spankalee4 days ago
Native compilation should be doable already with a Wasm compiler like Wastrel.

One reason I haven't explored that is that I want to tailor the language for the more constrained environment of Wasm GC first.

armchairhacker4 days ago
Ironically, since AI can build an STL and tooling (up to a full OS!), I think now presents an opportunity for a language that does start from scratch, at least targeting hobbyists (besides those specifically looking for novel languages).
demibabs4 days ago
A programming language for agents seems ill-conceived in my opinion.

Agents will naturally be bad at it due to a lack of examples.

dom964 days ago
Agents are fairly good even at languages designed to trick them. I built one[1] and it does make for a good benchmark[2] to see which LLMs are actually good. I think that a language which is largely similar to others will be a piece of cake for most and any advantage that an existing language will have will be minor enough to not matter.

Btw if folks have ideas of how to make Killswitch even harder for LLMs I’d appreciate them.

1 - https://killswitch-lang.org

2 - https://bench.killswitch-lang.org

spankalee4 days ago
From experience with Zena, this is not true at all. Opus, Fable, Gemini Flash and Pro all barely make any syntax mistakes after a little is in context, and those are caught extremely early.

The one thing I do see sometimes is that agents sometimes don't take advantage of added features, but that's partially because the Zena code base doesn't use them as much yet. I'm working on skills and linter-based suggestions to use better patterns.

ashton3144 days ago
This is absolutely not the case in my experience. I am building a very large embedded domain specific language for describing distributed systems. It looks like a small subset of Elixir, but with object-oriented syntax in a lot of places. (It’s called a choreography; there exist many other choreographic programming languages.)

Even though this programming language is absolutely nowhere in any large language model’s training set, they have so far done extremely well at extrapolating from the small set of examples I’ve given it when I need an agent to generate some tests or whatever for me.

whattheheckheck4 days ago
This keep getting repeated. So were just stuck with whatever we have at the point of training the magic plagiarism machine?

The future is cooked

ryuuseijin4 days ago
I think starting with a familar typescript-like base language is a good approach to this. This should be familiar enough for LLMs for the most part as long as additional features can be explained in a succinct system promopt/skill.
mqus4 days ago
I don't think any new languages will easily beat the existing ones, simply because of the mass of training data that is available. An LLM will have a much easier time one-shotting Java than even Rust today because it doesn't have to look up much and the ecosystem was stable for many years, so the internet is filled with content that is still up2date. Building a new language (or even altering existing ones) will take much longer to get into the models.

LLMs that have to constantly fix their code and look up libraries or features are much slower, more expensive and more prone to create inefficient and slightly wrong code. Sure, an LLM can just create a new web framework from scratch, but you usually don't want to spend your tokens on that.

I also think that with LLMs, new languages will have more trouble building an ecosystem, because until now, the popular libraries had a "proof of work" that kinda also made sure that they were well-trodden and maintained. Now, anyone can just push their new generated library and we have no clue on how to compare (and tbh reading LLM-generated READMEs is also not pleasant).

merelydev4 days ago
Agreed, said something similar yesterday:

"LLMs dont create anything new, if programmers stop reading the code technology will be forever frozen to 2022, no new programming languages, operating systems, concurrency primitives, databases, networking protocols, UI frameworks everything will be based on the training data and future generations will forget about all the primitives we now take for granted.

If someone creates a new programming language/ framework or new better way to do async or whatever, no one will use it because it is not in the training data and it wont take off because everyone is using LLMs. It will be like using the same Lego pieces over and over."

https://news.ycombinator.com/item?id=49854536

K0balt3 days ago
Mmmm idk. I used an LLM to write assembler in my made up API so I don’t think what you say is true. The value in LLMs is precisely that they are not just regurgitating training data, but rather inferring concepts extracted from trained data. If a programming language used concepts completely disconnected from existing paradigms you’re probably right… but that would also be quite challenging for humans to learn to use, since by intrinsic construction it would also be widely separated from human language.

Esoteric languages like BrainFuck are esoteric and difficult precisely because they go out of their way to eschew conceptual links to existing languages or paradigms.

So if you invented a new type of esoteric language with arbitrary syntax and strange operators (not sure how you’d do that, exactly, iirc all fundamental binary operators are known) it might be impossible to use with an LLM even if the user manual was in context… but aside from that, languages and the underlying concepts are extremely generalizable.

awb3 days ago
> I don't think any new languages will easily beat the existing ones, simply because of the mass of training data that is available.

Seems like a solvable problem though by generating synthetic data that’s guaranteed to be accurate through linters, compilers and tests.

If it’s ~9 figures to train a frontier model, it seems like training on a new language could be a rounding error if it was a priority.

I could see the appeal of a new agentic-friendly language that’s focused on minimizing tokens. The standard library could be massive with no concern of making the language easily readable or learnable.

merelydev3 days ago
> Seems like a solvable problem though by generating synthetic data that’s guaranteed to be accurate through linters, compilers and tests.

This makes sense and is very insightful, thank you. But with that solution it seems that only LLM companies will be in the position to create new languages.

fcarraldo3 days ago
Tremendous amounts of the Java out there in the training corpus would be based on outdated patterns, especially anything pre-Java 8 or pre-Java 21 (Pattern Matching, Project Loom). I don’t think this negatively impacts the quality of LLM Java code but it sort of calls into question whether this is simply a case of more is better.

My personal experience is that LLMs are quite good at one-shotting complex solutions in both Rust and Go, and tend to be idiomatic.

williamcotton3 days ago
I've been playing around with a new DSL, and I've had good luck adding a "describe" feature, eg,

  newdsl describe some_feature
This returns the docs for that language feature. So instead of a giant SKILL.md containing the language spec, it exposes a way for an LLM to introspect.

DSLs climb the ladder of abstraction and constrain the solution space at the language level, which IMO provides tighter feedback to both humans and machines.

reinhash4 days ago
The frontier models are exceptionally good at one shoting rust. I find them to be even better at writing rust than python (and the dataset for python must surely be bigger). I think there might be a threshold of enough data for a programming language to make it useful for an agente
jmull4 days ago
There's no point to try to adapt our languages to the strengths of LLMs when the strength of LLMs is working in terms of our languages.

Implement whatever abstractions you think LLMs should work in terms of in whatever language is handy, and have your LLM use those abstractions.

hollars4 days ago
I think I get your point but I think it's reductionist to the point of being incorrect. LLMs must be better at some semantics than others. Programming languages don't have random semantics, they have what matches the world and what matches our languages and so on. And the current frontier LLMs aren't so generic that they can predict any phrase no matter the quality of the content and grammar. Concretely I mean some languages are harder to reason about (predict) than others.
d_tr4 days ago
The strenghts and weaknesses of LLMs are changing, so starting a long-term project like a PL with them in mind sounds like a good way to end up with something obsolete before it is even usable.
jmull4 days ago
What semantics do you guess LLMs would better at, the semantics they are trained on, or something new with altered semantics?
YuechenLi4 days ago
> Coding agents don’t care about tedium.

LLMs are generally more tolerant of tedium than humans, but they make mistakes more often on repetitive mechanical tasks. I've had Claude write PTX directly once, and Claude just wrote majority of it and commented something along the line of "repeat this block 7 more time with these minor changes" instead of writing them out, so the code didn't work.

So, no, replacing compilers with LLMs is probably a worse option than having them code a compiler/programming language.

Plugging my own thing to use as example:

https://github.com/yuechen-li-dev/oct

Oct is the first programming language that I made with Codex. It started out as "Octave Modern" and was never intended to be a language for LLMs in the first place, but rather a teaching language that I've been thinking about for a decade because of my frustration with academic code and specifically reproducibility, with Python and Matlab in particular.

But it turns out the same design choices that was made to prevent bad patterns from academic code also made it pretty good for LLM coding: statically typed, immutable by default, GC'd with fast compile and runtime because it compiles to Go, along with features designed for scientific compute like SI units and builtin graphing etc.

But now it just took a life of its own, so it has extra features like templates/concepts, iterators, async/await, database query, build system for C/C++, SystemVerilog/WASM(WIP) codegen, LaTeX/pdf generation etc. None of them were features that were developed in isolation of "what an LLM agent might want to write" but to address a specific problem that I had encountered or to address specific failure modes that Codex/Claude actually had.

That's why I think Oct is probably one of the better languages for AI to write/generate, not because it was designed to be AI friendly in abstract, but that it's developed against how AI actually writes code, even though again, it is still very much a work in progress.

cmontella4 days ago
> I've been thinking about for a decade because of my frustration with academic code and specifically reproducibility, with Python and Matlab in particular.

I'm with you there, I think I starred Oct when I came across it. I've been on that quest since 2014, we should collaborate! I'll be presenting this work at IROS tomorrow, I'd love to hear what you think: https://mech-lang.org/iros-r4r-2026/index.html

YuechenLi4 days ago
Yeah, we definitely should collaborate on something, since I think we both came to the same conclusion that explicit state machines should be the primitives of a programming language.

So here is my experiment with Kalman filters here.

https://github.com/yuechen-li-dev/oct/tree/main/Experiments/...

It's been a while, but I think the findings there are mostly that Kalman filters are relatively heavy and fairly narrow in application in that it's good at filter Gaussian noise out but not much else, and it's kind of branchy so it doesn't run on the GPU very well, you can try to see if a simpler feedforward like Smith predictor can work as well.

Also, something to try out: the explicit state machine stacks/pushdown automata are only half of the equation, the other (and imo more important) half is argmax/utility AI based transition instead of traditional state machine graph.

But yeah, if your target is embedded/bare-metal application for robotics, since Oct really isn't designed for it, maybe you would like to check out what I'm currently working on, the Concept programming language?

https://github.com/yuechen-li-dev/Concept/

adastra224 days ago
That's a nice language, btw.
YuechenLi4 days ago
Thanks. The fun thing about the name "Oct" is how many dumb puns I can make with it. For example, the LaTeX/PDF generation functionality is called Oct-cument.

I'm not proud of that pun.

Animats4 days ago
LLMs are good at optimizing towards local goals. Getting types right at compile time is a local goal. Entry and exit assertions are local goals. Unit tests are local goals. So those constructs all help AI-generated code.

Matching a desired output is a global goal, but even that sometimes works now. Someone sent me a LLM-generated JPEG 2000 decoder. They got Fable to generate a decoder that uses a GPU to get the same answer as the reference implementation gets on the GPU.

te_chris4 days ago
I’ve reimplemented some python science code into rust using LLMs and it absolutely works well if you setup the criteria for done. In my case it was GIS-ish code so a key requirement was identical in/out pairs at 10k points across CONUS. It worked, and when it didn’t it showed up bugs that the agent would go and research, read up on other implementations, then try again.

Net result faster (10-70x) and massively lower memory footprint. Now we’ve got a service that can do something essentially instantly that most people wouldn’t have even bothered to try before.

Read the full thread on Hacker News →

Related stories