AI Already Stole the Developer’s Job

Aug 17, 2026·
Felipe Cardoso
Felipe Cardoso
· 9 min read
blog

Hard words to swallow: AI is not going to steal the developer’s job. It already stole it.

Every week another article asks whether AI will replace programmers. The premise is wrong. It is not a question of when. It already did.

The job that existed in 2022 does not exist anymore. What remains has another name, another nature, and another skill set. The name is harness engineering.

Anyone still debating whether to “adopt AI” is debating whether to board a train that has already left.

This is my opinion, formed by running this workflow every day, not a comfortable prediction about some distant future.

How I work now

I am a backend and infrastructure engineer. I alone own the AWS infrastructure at my company, built from scratch with CDK and a lot of AI harnessing, and I carry production on-call.

When something breaks at 3 a.m., my phone rings. I am not theorising from the outside. This is daily lived experience.

I now barely write code. I make the strategic and high-level architectural decisions, describe them in natural language, and AI agents implement.

I almost never read diffs. I validate behaviour through unit and end-to-end tests, which are also driven by AI.

You may be thinking, “this guy no longer knows how to program.” Wrong. Reviewing code stopped being the highest-leverage use of my attention.

AI reviews the mechanical part better and faster than I do. Programming became a commodity. It always was. The price tag has finally become visible.

Human review still exists, but it moved to the expensive places: irreversible changes, architecture, threat models, IAM, cost, and defining what a test proves.

So if I do not write the system, what do I build? That is the right question. I build the harness.

Building and fixing became too cheap

The central thesis is simple: the cost of building and fixing software collapsed.

When regenerating an implementation costs close to zero, reviewing every line of that implementation becomes a micro-optimisation of a resource that is no longer scarce.

This resembles the transition away from reviewing compiler-generated assembly. The comparison has a limit: compilers have stable semantics and decades of maturity. Agents are probabilistic.

That is why agents do not deserve blind trust. They deserve a system that proves, limits, and reverses their output.

Bugs, leaks, state corruption, and exploding cloud bills existed before LLMs. The question was never how to prevent every error. It is how quickly we detect, contain, and fix one.

Answering that automatically, continuously, and cheaply is exactly the human job that remains.

Harness is the new code

If code is disposable, all of your trust has to live somewhere else. That place is the harness.

The term comes from a test harness: the structure that holds a component in place so you can stress it without letting it fly across the room.

For agent-built systems, it means the entire apparatus that continually answers one question: does this work inside the limits I accept?

The first time I saw AI harness articulated clearly was in Anthropic’s post on harness design for long-running autonomous development.

Their point is not to wrap the model in a larger prompt. It is to build a planner, generator, and evaluator around contracts, state artifacts, and a feedback loop.

The important detail is separating the generator from the judge. An agent reviewing its own work tends to approve itself. A sceptical evaluator with criteria and access to the running system finds the fake feature the generator calls done.

In practice, my harness includes:

  • Specifications decomposed into small tasks, behaviour contracts, and observable acceptance criteria
  • Unit, integration, regression, and end-to-end tests that validate behaviour rather than implementation detail
  • Human-written oracles, invariants, and adversarial examples, because AI-generated tests without an oracle are theatre
  • A separate evaluator agent with access to environments and traces, able to test as a user and as an attacker
  • Deterministic gates: linting, type checking, tests, builds, SAST, dependency policy, and infrastructure validation
  • Structured logs, metrics, traces, actionable alerts, and correlation across deploys, agent runs, and changes
  • Feature flags, canaries, disposable environments, cheap rollback, and an explicit blast-radius limit
  • Cost caps, least-privilege policies, drift detection, and regular configuration audits

The harness separates vibe coding from shipping slop. Developers stopped developing systems and started developing the harness that lets systems be generated.

Code became output. The harness became the product. That is AI harness engineering, and it is the foundation of the profession from here on.

The failure ledger is the memory that is missing

The trick is making the harness feed itself. Every meaningful production failure must become an artifact that changes the next round’s behaviour.

I call that a failure ledger: a record of everything that broke and what was done to prevent or detect its repetition.

It is not a postmortem that dies in Notion. It is operational data. Each entry links incident, hypothesis, signal, cause, correction, test, alert, version, and owner.

It is not RAG either. RAG stores facts about the world. A ledger stores what the agent and the system did in the world, the context of that decision, and the observed consequence.

Daice Labs makes the core argument clearly.

What an AI cannot remember about its outcomes, it cannot fix

Without causal history, the next incident is inference rather than evidence.

A useful entry has, at minimum, an incident ID, commit and deploy image, prompt and model version, tool calls, sanitised inputs, broken assertions, trace, metric, blast radius, and the countermeasure created.

The countermeasure is not just “fixed the bug.” It can be a regression test, a database invariant, an IAM policy, a cost alarm, a canary, or approval before an irreversible state change.

The ledger should be append-only. A correction is a new event, not an edit to history. Without that, you cannot answer what the agent decided, why it decided, or which earlier data contaminated the decision.

My goal is simple: the same failure should not teach us the same lesson twice. Dependencies change and tests remain incomplete. But forgetting what broke is a choice, not fate.

Real robustness does not come from theory badges. It comes from a system that got hit by production and recorded every hit.

How Netflix uses Chaos Monkey to expose controlled failures and prove automatic recovery, the same principle an AI harness applies to agents

Inject controlled failures into the agent and deploy before production injects them for you.

For agents, test tool timeouts, 429 responses, stale schemas, out-of-order queues, denied IAM permissions, partial migrations, lying caches, and prompt injection in tool output.

Do not merely test whether the agent completes the happy path. It must survive bad data, slow dependencies, revoked authorisation, and an external effect that reports success without happening.

The ledger only learns from failures that happen and are detected. It is blind to silent corruption, excessive permissions nobody has abused yet, and slowly leaking cost.

The response is more harness: billing anomaly detection, data reconciliation, schema-drift checks, secret scanning, regular permission audits, and blast-radius limits.

Outsource the sensitive parts, vibecode the rest

“But what about critical parts? Auth, payments, identity?” You should not be hand-writing those in 2026, with or without AI.

Stripe, Okta, Clerk. Companies with whole teams dedicated to these concerns do them better than you ever will. Outsourcing sensitive capabilities is good engineering, full stop.

What remains is integration, configuration, and your domain data. That is where most failures I see live: a public S3 bucket, permissive IAM, an unsigned webhook, or bad glue between components.

The error surface moved from code to glue. In other words, it moved exactly into harness territory.

  1. Buy the sensitive parts: auth, payments, and identity.
  2. Vibecode the disposable parts: almost everything else.
  3. Put the human brain into the harness: tests, observability, configuration, and risk decisions.

That is not laziness. It is rational allocation of the system’s most expensive resource: your attention.

CRUD developers are already done

Rational allocation has a darker side: it reveals who was allocated in the wrong place.

Think of the old-school developer: receives a ticket, writes an endpoint, builds CRUD on an ORM, maps request to query to response, and moves from sprint to sprint.

Internal backend, form to database, report to screen. Stable life.

That work was first to become a commodity for a reason. It is the most predictable, repetitive code and the best represented in training data.

Near-zero ambiguity. An agent delivers it in minutes for cents, with tests. There is no defensible scenario in which hand-typing it is a good use of an engineer’s salary.

The uncomfortable part is that this developer is not “going to be” replaced. They already were. The badge still exists; the function does not.

The role survives through organisational inertia, and inertia is a deadline, not protection. When a company learns that one agent operator delivers the CRUD backlog of N developers, headcount math resolves itself.

This is extinction with delay, and the delay is getting shorter.

Learning to write a CRUD does not automatically teach someone what to measure, when to distrust a result, or how to design a test that proves something.

The escape route exists, but it is not “learn prompting.” Prompting is trivial. It is learning to build harnesses.

The developer became an operator

Put all of this together and the picture is obvious. The developer using AI today is no longer a developer in the 2022 sense. They are an operator.

People who do not understand system design, product, security, and harnesses are already behind, whether they know it or not.

Typing code became the punch card of our era. What did not become legacy is the foundation: knowing what to build, how to structure it, what to measure, and when to distrust it.

The value layer moved up. People clinging to the lower layer are competing with an API.

There is a side effect few people discuss: the pyramid flattened. Juniors who only wrote code lost their function, yet the path toward architectural judgement used to run through years of writing code.

The industry has not solved where the next operators will come from. That makes those already across the bridge scarcer, not less scarce.

The lag nobody tells you about

The job changed, but the market did not. Most companies still hire, interview, and pay for the old model: LeetCode, live diff review, system design without AI.

There is a brutal lag between what the job became and what the hiring funnel measures.

The practical answer is two modes. Operate the new way every day, and keep interview mode warm in parallel.

Annoying? Yes. It is the toll for being ahead of the curve while the market catches up.

Conclusion

The “will AI replace developers?” debate already died. Nobody told everyone.

Replacement happened in the nature of the work, not in the job listings. Writing code became orchestrating agents with judgement. Code review became harness engineering. Knowing the codebase became knowing the system through its signals.

The bet is not to defend the old skill. It is to become excellent at what remains human in the loop: decision, architecture, harness, and final responsibility for what runs.

Building and fixing became cheap. Judgement is still expensive. Charge for it.

Felipe Cardoso
Authors
Senior Backend Engineer
Backend systems, cloud infrastructure and distributed systems. Currently in Tokyo.
Loading comments…