Emerging Frontiers · 6 min
Self-Improving Systems: A Research Question, Not a Product
Break the phrase apart. Each mechanism people file under "self-improving" — memory, prompt optimization, code modification, fine-tuning from experience — has its own maturity level and its own safety profile, and only some of them are safe to deploy.
Ship a model that rewrites its own prompts overnight and by morning you have a system nobody on your team designed. That is the whole pitch of "self-improving AI," and it is also the whole problem. The phrase sells a product that does not exist yet and waves away a research question that is genuinely open. When someone tells you their agent "learns from experience" or "improves itself," the useful response is not awe. It is a question: which of about nine different things do you mean, and which one did you actually build?
"Self-improvement" is not one capability. It is a bag of loosely related mechanisms that got stuffed under one label by marketing. Each sits at a different maturity level, and each has a completely different safety profile. Pull them apart and the hype collapses into a set of concrete engineering choices: some boring and shipped, some speculative and dangerous.
Take the phrase apart
Here is the decomposition. Nine mechanisms, roughly ordered from "you already do this" to "nobody does this reliably":
1. Better context from previous runs — stuff last session's summary into the prompt 2. Persistent memory — a vector store or file the agent reads/writes 3. Prompt / workflow optimization — search over prompts against a metric (DSPy) 4. Synthetic-data generation — the model writes its own training examples 5. Automatic evaluation — the model grades its own or others' outputs 6. Code modification — the agent edits the code it runs on 7. Test-time parameter updates — weights change during inference, per task 8. Fine-tuning from accumulated experience — retrain on logged interactions 9. Autonomous experiment / research loops — propose, run, and act on its own experiments
Items 1 and 2 are shipped, mundane, and mostly a retrieval problem. Item 3 is real and reproducible; I'll come back to it. Items 4 and 5 work but share a nasty failure mode. Items 6 through 9 are where the interesting research and the interesting hazards live, and where most product claims run wildly ahead of the evidence.
These do not compose cleanly. A system that generates its own training data (4), grades it with itself (5), and then fine-tunes on the result (8) is a closed loop with no external ground truth. That loop is exactly where model collapse shows up: train repeatedly on your own outputs and the distribution narrows, the tails vanish, and quality degrades in ways that stay invisible until they are catastrophic. Self-reference is not automatically improvement. Often it is just drift with a good story.
The one that actually works: prompt optimization
If you want a concrete, defensible example of self-improvement today, it is automatic prompt and pipeline optimization. DSPy (Khattab et al., 2023) is the cleanest instance. You stop writing prompt strings by hand and instead declare a program — modules with typed input/output signatures — plus a metric and a training set. An optimizer then searches over prompt phrasings and few-shot demonstrations to maximize the metric.
# Sketch of the DSPy pattern, not runnable verbatim
class RAG(dspy.Module):
def forward(self, question):
ctx = self.retrieve(question).passages
return self.generate(context=ctx, question=question)
optimized = optimizer.compile(RAG(), trainset=train, metric=exact_match)The system "improves itself" in a bounded, auditable sense. The search space is prompts. The objective is a metric you chose. The output is a diff you can read. Nothing about the model's weights or its runtime changes. This is the safe end of the spectrum precisely because every part of the loop is external and inspectable: you own the metric, you own the trainset, and you can diff the result against the previous version before you ship it.
That property — can I read the diff? — is the line that separates the mechanisms you can deploy from the ones that are still research questions.
Different mechanisms, different threats
Now walk the same list through the lens of the security modules — poisoning, supply chain, provenance — and the profiles diverge sharply.
Persistent memory (item 2) turns every past interaction into part of a future prompt. That is a poisoning surface. If an attacker can write to the memory store — through a tool call, a shared document, a prior conversation — they can plant an instruction that fires days later, in a different session, against a different user. This is indirect prompt injection with a delay fuse. The OWASP Agentic AI threat work names memory poisoning as a first-class agentic risk: memory and tool state are attack surfaces, and "the agent remembers things" means "the agent can be taught things by whoever can write to its memory."
Code modification (item 6) is the one that should stop you cold. An agent that edits and re-runs its own code has, by construction, defeated your review process. The point of code review and signed releases is that a human or a pipeline gate sits between "code changes" and "code runs." Self-modification removes the gate. Provenance breaks with it: you can no longer answer "what code produced this output" with a git SHA, because the code mutated between commit and execution. Every guarantee you built on the supply chain — pinned dependencies, reproducible builds, an audit trail — assumes the artifact you reviewed is the artifact that ran. Self-modifying code violates that assumption on purpose.
Fine-tuning from experience (item 8) inherits the poisoning problem and makes it permanent. A prompt injection lives for one session. A poisoned training example baked into weights lives until you retrain, and you may never know it is there. If your logged interactions become training data with no provenance check, then anyone who can influence a logged interaction can influence your next model. The disciplined version demands what you would apply to any training corpus: know where each example came from, and never treat "the model produced it" as a proxy for "it is correct."
One question ranks all nine. After the system acts, can you still reconstruct what happened and why? Prompt optimization, yes — there is a diff. Memory, mostly, if you log every write with its source. Code modification and experience fine-tuning, not without deliberate infrastructure you probably have not built yet.
What to actually do with this
Treat "self-improving" as a claim to be decomposed, not accepted. When you hit the phrase in a vendor deck, a paper abstract, or your own design doc, run it through three questions:
1. WHICH mechanism? Map it to the nine. "Learns from experience" almost always means item 1 or 2 (context/memory), not 8 (fine-tuning). 2. What is the ground truth? If the loop grades itself with itself (4 + 5 + 8), there is none — expect drift, not improvement. 3. Can you read the diff? If the change is inspectable (a prompt, a demonstration set), it is deployable. If it mutates code, weights, or memory with no audit trail, it is a research project.
Then build for the mechanism you actually chose. If it is memory, treat the store as untrusted input: validate on read, log every write with its source, and never let memory silently override system instructions. If it is prompt optimization, version the optimized artifacts like code and diff them before promotion. If it touches weights or executable code, keep the human gate — the model proposes, a reviewed pipeline disposes.
The honest state of the field: most of "self-improvement" is either mundane (retrieval), bounded and real (prompt search), or genuinely unsolved (open-ended self-modification that stays safe and correct). None of it is a shipped capability you can buy and trust to improve unattended. The people doing serious work here say so plainly; the people selling it do not. Knowing which mechanism is on the table, and what it does to your provenance story, is how you tell them apart.
Sources
- Khattab, O. et al., "DSPy: Compiling Declarative Language Model Calls into Self-Improving Pipelines," 2023. https://arxiv.org/abs/2310.03714
- Shumailov, I. et al., "The Curse of Recursion: Training on Generated Data Makes Models Forget," 2023. https://arxiv.org/abs/2305.17493
- OWASP GenAI Security Project, "Agentic AI – Threats and Mitigations," 2025. https://genai.owasp.org/resource/agentic-ai-threats-and-mitigations/
- OWASP GenAI Security Project, "OWASP Top 10 for LLM Applications 2025," 2025. https://genai.owasp.org/llm-top-10/
- Yao, S. et al., "ReAct: Synergizing Reasoning and Acting in Language Models," 2022. https://arxiv.org/abs/2210.03629
- Madaan, A. et al., "Self-Refine: Iterative Refinement with Self-Feedback," 2023. https://arxiv.org/abs/2303.17651