The Case for an AI Programming Model
TLDR: AI coding agents are tools, not people. Just because we can communicate with them in natural language doesn’t mean we should anthropomorphize them. Instead, programmers need a clear mental model of their behavior, which requires a programming model with constraints and guarantees.
Many people ask me how I come up with research ideas, and while it’s different for everyone, my favorite approach is to: 1) think about why something is harder to use than it should be and 2) build a system to make it easier. For example, my most recent work has been on making kernel-bypass hardware easier to use and my PhD thesis focused on making mobile-cloud applications easier to build. For the last year, I’ve been thinking hard about AI coding agents, which presented a perfect target for my preferred type of research brainstorming.
AI coding seems so promising. It lets novice programmers interact with a human language API without learning programming languages. It lets advanced programmers rapidly prototype and build large systems. And yet, it is falling far short of that promise. Novice programmers are not learning any skills that they need to move beyond building small bits of code. Advanced programmers are causing catastrophic failures and leaving mountains of technical debt. This blog post explores this gap between promise and reality and what we need to do to fix it.
The Problem
AI coding agents are the biggest programming advancement in the last decade. They let programmers work at a higher level of abstraction than existing programming tools. Instead of using a programming language, developers can describe software using natural language and then coding agents turn that description into code. This advance has been compared to the industrial revolution in terms of moving people beyond manually crafting software.
Unlike other higher-level abstractions in computer science, coding agents have no programming model. Take consistency models, which serve as abstractions over complex storage systems. They provide guarantees: for a given set of inputs, the consistency model (e.g., linearizability) constrains what the storage system may output. They hide complexity: programmers do not need to understand the implementation of a linearizable storage system to use it. And most importantly, they let programmers construct a mental model: a linearizable storage system is suppose to represent an idealized memory store.
AI coding agents provide no guarantees. There are no constraints on the inputs (i.e., prompts) or guarantees on the output. Indeed for the same input, any AI agent will return different results each time. It often takes multiple tries before a programmer can get a coding agent to return the correct code. There is no hiding of complexity. For each try, the programmer must review the outputted code, which is often verbose or incorrect.
Worst of all, there is no way to develop a mental model of the complexity that AI coding agents abstract. Programmers often will develop an ad-hoc mental model because humans are excellent at recognizing patterns in the chaos; however, these models are generally incorrect and break as soon as a new frontier model appears. There is no way to teach students to us AI because there is nothing to teach. The entire process is guesswork, and a monkey hitting a key repeatedly might have similar success rates.
The absolute worst trend is the personification of AI coding agents. Just because we can use human language to communicate with them does not make them junior engineers. Humans love to anthropomorphize things. My coffee maker has googly eyes on it; that doesn’t make it a barista. AI coding agents are tools and as such, we must have the same expectations as all other programming tools. They must be useful, reliable and have a clear set of guarantees. Without that, we are simply wasting our time staring into the clouds and expecting shapes to appear.
The Solution
AI coding agents most closely resemble compilers and assemblers. They take a higher level language and return a lower level representation. For example, a C compiler takes C code and returns assembly. The ISO C language specification has a clear set of rules for the compiler, and it is possible to use any C compiler to compile C code without needing to check the assembly each time. While there are exceptions, like undefined behavior, these properties are generally true for all C compilers and other language compilers.
My goal is to turn AI coding agents into a more general-purpose, usable programming tool, similar to compilers and assemblers. More specifically, I want to build a system (or set of systems) that will wrap current AI coding agents and enforce a set of compiler-like guarantees. AI coding agents provide messy, unreliable code so it is up to researchers to correct and fix that output to something usable for programmers. I do not know exactly what those guarantees or systems will look like but I propose the following properties (there might be more!):
- Well-defined behavior. Any under-specified or vague prompts should be corrected or thrown away. For the set of acceptable inputs, there should be a guaranteed set of properties provided by the outputs.
- Deterministic. The same output is returned every time for the same prompt.
- Model-oblivious. Different coding agents using AI models are indistinguishable (except for performance).
Obviously, defining the exact programming model will require some thought and experimentation, but some starting points for research directions include:
-
How do we constrain the inputs? Do we need a more formal natural language? Do we throw away natural language altogether and use a more algorithmic, higher-level programming language?
-
How do we filter for incorrect outputs? Can we build compilers that detect these and return errors to the AI agent? Can we build more tools that provide input to the AI agent and eventually lead to deterministic outputs?
-
Do we do all of this at compile time or do some of these things need to be enforced by a runtime system (or runtime code inserted by the compiler)?
If you are interested in working in this area with me, contact me. I’m taking interns for the next summer!