The 27-Year-Old Bug: How Can a Vulnerability Hide in OpenBSD for Decades?

D. Rose · 18 August 2026 · 5 min

In 2026, Anthropic said Claude Mythos Preview found a vulnerability in OpenBSD that had existed for roughly 27 years. It also found a 16-year-old FFmpeg bug in code that automated tests had exercised millions of times without recognizing the security problem.

In 2026, Anthropic said Claude Mythos Preview found a vulnerability in OpenBSD that had existed for roughly 27 years. It also found a 16-year-old FFmpeg bug in code that automated tests had exercised millions of times without recognizing the security problem.

That sounds impossible until you understand what security testing actually sees.


The 30-Second Version

Software can execute a line of vulnerable code millions of times without triggering the specific state required for exploitation.

Think:

code executed
      ≠
bug triggered
      ≠
bug recognized
      ≠
security impact understood

A vulnerability can survive because:

  • the input combination is rare,
  • the bug spans several functions,
  • tests don't assert the right property,
  • fuzzers cannot reach the necessary state,
  • the crash looks uninteresting,
  • reviewers misunderstand an invariant,
  • no one asks the exact security question.

AI may add a different search strategy: semantic reasoning over code and behavior rather than only generating random inputs.


Part 1: “But OpenBSD Is Security-Focused”

Exactly.

That is why the finding is interesting.

Security-hardened does not mean bug-free.

It means the project invests heavily in things like:

secure defaults
code review
memory safety mitigations
attack-surface reduction
proactive auditing

But a large operating system contains enormous amounts of stateful code written over decades.

One forgotten assumption can survive.


Part 2: WTF Is Fuzzing?

A fuzzer repeatedly throws inputs at software and watches for bad behavior.

Input 1 → okay
Input 2 → okay
Input 3 → crash
              ↓
        interesting!

Modern fuzzers are smarter than pure randomness.

They mutate inputs based on which code paths they reach.

Their goal is often:

maximize code coverage
+
trigger crashes / sanitizer failures

Fuzzing is incredibly effective.

But coverage is not comprehension.


Part 3: Code Coverage Is Not Security Coverage

Suppose a test executes:

if (length < buffer_size) {
    copy(data, length);
}

one million times with safe lengths.

The line is covered.

The dangerous boundary condition may never occur.

So:

line executed 5,000,000 times

does not prove:

all meaningful states tested

This is one reason old bugs survive.


Part 4: State Space Explodes

Imagine a function depends on five variables.

Each has 100 possible values.

The theoretical state space is:

100^5 = 10,000,000,000 states

Now add timing.

Concurrency.

Network history.

Allocator layout.

Previous packets.

Configuration.

Permissions.

Software state grows combinatorially.

Testing everything is impossible.


Part 5: The Difference Between “Crash” and “Exploit”

A fuzzer may find a crash.

A security researcher asks:

Can attacker control it?
Can it write memory?
Can it leak information?
Can mitigations be bypassed?
Can it cross a privilege boundary?

Those are reasoning questions.

A model that understands code structure may help connect:

weird behavior
violated invariant
attacker-controlled state
security impact

That is qualitatively different from “generate more random inputs.”


Part 6: Why the FFmpeg Example Is So Good

Anthropic said Mythos found a roughly 16-year-old FFmpeg vulnerability in a line of code that automated testing had hit about five million times.

That is a perfect illustration.

The question is not:

“Did fuzzing touch the line?”

The question is:

“Did fuzzing construct the precise semantic conditions under which touching the line becomes dangerous?”

Those are not the same thing.


Part 7: AI Is Not Magic Static Analysis

Traditional static analyzers inspect code without running it.

They look for patterns:

use-after-free
unchecked bounds
uninitialized data
unsafe API

They are excellent at specific classes of defects.

LLMs potentially add something different:

read surrounding implementation
infer developer intent
identify suspicious assumption
construct hypothesis
test hypothesis dynamically

That resembles a human researcher.

The model does not replace static analysis or fuzzing.

It can orchestrate them and reason between them.


Part 8: The “Forgotten Invariant” Problem

Security bugs often come from an assumption like:

"This value can never be negative."

Twenty years later a new code path makes it negative.

No one updates the old function.

Now:

old assumption
+
new caller
=
new vulnerability

The vulnerable line may be ancient even if the exploitability is newer.

That is another reason age alone can mislead.


Part 9: Why Humans Miss Old Bugs

Humans have cognitive shortcuts.

If code has been stable for 20 years, reviewers assume:

“Someone would have found a serious bug by now.”

That creates a weird form of inherited trust.

The longer code survives, the safer it feels.

But age can also mean:

old assumptions
old language
old parser
complex compatibility logic
few remaining experts

Exactly the things worth reviewing.


Part 10: Why AI May Be Good at Archaeology

Models can cheaply read enormous volumes of old code without boredom.

They can ask the same annoying questions repeatedly:

Who controls this value?
What happens at the boundary?
Who validates this pointer?
What invariant is assumed?
Can this caller violate it?

That is software archaeology at scale.

A human team rarely has budget to deeply re-audit every 1999-era subsystem.

An agent can.


Part 11: But False Positives Still Matter

Models can hallucinate vulnerabilities.

They may misunderstand build flags, unreachable code, runtime invariants, or mitigations.

So the responsible workflow is:

AI hypothesis
reproduce
minimal test case
confirm root cause
assess impact
human review

AI lowers search cost.

It does not eliminate verification.


Part 12: What This Means for Legacy Software

The scary part is not one OpenBSD bug.

It is the amount of old code running critical infrastructure:

network stacks
media parsers
compression libraries
database engines
industrial protocols
kernel drivers
crypto glue code

If AI makes deep auditing cheap, we may discover that “mature” software contains a much larger latent vulnerability inventory than expected.

That is bad in the short term.

Potentially excellent in the long term — if defenders patch first.


The Big Misconceptions

“Millions of fuzz executions should have found it.”

Only if the fuzzer reached the right semantic state and recognized the consequence.

“Old code must be secure because attackers have looked at it for years.”

Age is evidence of exposure, not proof of correctness.

“AI discovered the bug just by reading English-like code.”

Frontier vulnerability workflows combine reasoning, tools, execution, and validation.

“This means fuzzing is obsolete.”

No. Fuzzing remains one of the strongest automated security techniques; AI can complement it.


If You Remember Only Five Things

  1. Code coverage is not state-space coverage.
  2. A line can execute millions of times without the dangerous condition occurring.
  3. Security research requires connecting behavior to attacker-controlled impact.
  4. Old code accumulates trust and forgotten assumptions.
  5. AI's advantage may be relentless semantic code archaeology, not magical bug detection.

Sources & Further Reading