This is a synthesis of ideas I’ve overheard, read, discussed, and learned first-hand. It draws on over three years exploring LLMs for shaping code. My initial notes for this were part of an internal workshop I hosted for our engineering team at TORTUS Health, where we use and improve these techniques every day as we work to eliminate error in medicine.
The shifting bottlenecks
With the explosion in LLM capabilities, producing syntax got cheap. But good judgement did not. In many ways, AI makes the easy part easier and the hard part harder.
The bottleneck has shifted. It is no longer typing code. The new bottlenecks (for now) include:
- Deciding what to build.
- Verifying that you built it right.
- Understanding the stack, domain, and tools well enough to keep doing both.
Agentic engineering is, therefore, using AI to increase leverage, whilst preserving human ownership. That ownership enables safety, engineering quality, and deep understanding.
Caveat coder
Agents magnify our execution. But they do so indiscriminately for both the good and the bad. It is our discipline, norms, and systems which determine the balance. Be careful that, in the pursuit of becoming a 10x engineer, you do not become a 0.1x engineer. Or worse.
Automation applied to an efficient operation will magnify the efficiency. Automation applied to an inefficient operation will magnify the inefficiency.
– Bill Gates (apparently)
And as Zvi observes, AI helps you work. It also helps you avoid work, or merely look busy.
So know your mode…
Vibe to learn. Engineer to ship. Natty when it really matters.
I’ve written before about the differences and applications of vibecoding and agentic engineering. But here is how I’ve been thinking about the modes of software engineering in 2026:
- Vibed: built to learn. Explore an idea through a quick, disposable prototype. The code can go unreviewed if you only keep the lessons.
- Engineered: built to ship. AI-assisted, with a human owner aligning it to requirements and standards. Production-shaped: author-reviewed, tested, evidenced. Ask: is this safe to deploy at scale to real users? If it breaks, can I figure out why and how to fix it?
- Natty: unenhanced, built by hand, from first principles. Use this when the exact judgement, wording, argument, algorithm, or understanding is the point. Increasingly rare, but increasingly undervalued.
Take care to recognise that these modes are not a moral hierarchy. Vibing is only bad when it’s lazy or misrepresented as production-ready. Natty is not anti-AI, but simply reserves direct human craft for the crux of the problem.
Agents are fast at producing the median, but deep insight is never in the median. Writing is thinking and explaining a thing proves you understand it. For instance, preparing these notes for a workshop took me a very long time. So did adapting them for this essay. I could have one-shotted a document of equivalent length with an agent, but writing and structuring it was the work, not the production of the artifact.
I increasingly find myself working either with several parallel agents in a frenetic hive-mind or sitting quietly with a notebook and favourite pen. Taleb’s barbell strategy is hard to beat and, sometimes, you’ve just got to go fully natty!
It is good practice to name the mode before asking for feedback. A vibe check is not a meticulous PR review. This will be a defining characteristic that separates dysfunctional teams from great ones in 2026 and possibly beyond.
Fight entropy. Be cognitively kind.
Code is abundant. Our attention, time, and energy are scarce. Three anti-patterns consume them:
- Workslop: low-information output that looks decent when skimmed but says little when read. Poor work masquerading as good work. Slack messages, slides, bloated docs, and pretty much everything on LinkedIn.
- Slop grenades: unreviewed, unrefined generated output lobbed at colleagues. A workslop artifact that was quick to make but tedious to decipher. Code, text, slides, whatever. Your shortcut becomes someone else’s verbosity explosion with a side of cognitive shrapnel.
- Context rot: declining reliability as context grows. Stale documentation compounds the problem. A mistake in an AI meeting transcript goes uncorrected in the AI summary. The summary becomes an AI plan that AI executes. Qualifications disappear and errors survive. To expand the aphorism, it’s: a Xerox of a Xerox of a broken telephone of a hallucinated half-truth misremembered.
Principles of Agentic Engineering
You own the work. A human remains accountable for correctness, risk, cost, customer impact, and downstream consequences. “The agent wrote it” explains the mechanism. You still own the result.
Laziness is discipline. LLMs lack the virtue of laziness, as Bryan Cantrill argues. They don’t feel the future maintenance cost of another abstraction or another thousand lines. Left unchecked, agents simply make systems larger. They add 300 lines of comments and unit tests for a one-line fix. You must supply the discipline that makes them better.
Agents reward fundamentals. Being a top-tier engineer has never been more important. Good docs, correct types, discerning tests, robust tools, reliable infrastructure, and clear goals. The good and bad patterns in pull requests are the same as before. Only more so.
Okay is now cheap. As Ben Wallace puts it, craft, focus, and expertise are your moats. Make work you can be proud of.
Reality is the truest reviewer. Generated code is harder to review because you didn’t build the context by writing it. Building software is learning. Sometimes, implementing the crux yourself is how you acquire the understanding needed to judge it.
Build so you can collide with reality quickly and safely. Use tests, shadow runs, staging, logs, and real feedback. All the review bots in the world cannot replace running the thing.
LLMs feel easy but they aren’t. Wielding them well requires a deep model of what they can and cannot do. Simon Willison’s weird, overconfident intern is a useful analogy: imagine someone off the street who eagerly reads all your documentation, then does something absurd.
A model can write excellent Rust, then recommend walking to the carwash because it’s only 50 metres away. This is jagged intelligence, where capabilities spike around verifiable domains, but fail spectacularly elsewhere. Develop your intuition for those peaks and troughs. Keep updating it as models change. Optimistic skepticism is the right attitude.
Guardrails aren’t guarantees. LLMs can follow instructions embedded in content, including instructions that aren’t yours. Beware the lethal trifecta of private data, untrusted content, and external communication. An agent that reads incoming emails, queries customer records, and sends replies has all three. An attacker can hide instructions in an email and try to extract your private data. By default, pick at most two and enforce restrictions through permissions and isolation. This applies to the software you build and the tools you give yourself.
CLI over MCP. Where possible, give agents command-line tools. You and the agent then use the same tools and can help each other to use them even better. Agents are good at reading
--help, discovering commands, and controlling output. Once you’ve found the command that works, capture it in a script. Some MCP configurations load tens of thousands of tool-definition tokens upfront and make it absurdly easy to mix private data and external communication, which can get you into the lethal trifecta.Compound the capabilities. In compound engineering, each piece of work should make subsequent work easier. Mistakes become tests and warnings. Decisions and principles become docs. Repeated workflows become skills. Repeatedly chained skills become automations. Keep them under version control and review them with the code they influence.
Tacit tips & tricks
The above principles should outlive this year’s model improvements. However, practical fluency comes only from playing, noticing patterns, and developing intuitions of your own. Such tacit knowledge cannot be taught, but here are some ways you can help yourself in acquiring it:
Git is your guardian angel.
diff,status,reset, branches, and worktrees defend against slop and overwhelm. Good news: we’re all git chads now. Ask the agent to write the tricky commands, then inspect them. Master the primitive, and the agent helps you wield it.Curate your context. Supply the goal, constraints, examples, relevant state, docs, and definition of done. Context is the model’s attention in both the literal and metaphorical sense. Work with your harness’s instruction hierarchy. Learn how it prioritises system instructions, user prompts, skills, and repository docs. Explore broadly early: code, docs, conversations, git history. Then decide and distil. Compact tactically, and check the resulting summary. Correct mistakes before they become the next session’s assumptions. Limit noisy tool output with options such as
--quiet. Use subagents for bounded exploration, bringing back the findings that matter.Refresh the sesh. Compaction carries a summary forward. A fresh session lets you drop the previous framing and seek a cheap second opinion. You ask the builder: “Is this PR ready?” “Certainly! Amazing work. You’re the greatest engineer.” Then you open another session: “Grill the shit out of this thing.” “Oh, you’re so right. Look at all these gaping flaws.” Both answers can be sycophancy. Ask for concrete failures, reproduction steps, and evidence. Change the task from defending the feature to finding where it breaks.
Meta-prompting helps you separate exploration from execution. Under this pattern try starting with one agent surfacing trade-offs and edge cases. You turn those findings into a spec, then hand it to another agent to execute.
Discuss, decide, delegate, delete. Work in 4 discrete phases: 1) Explore and align on the problem. 2) Lock the spec. 3) Hand the bounded task to the agent. Then 4) cut bloat, dead code, and creeping scope. The final deletion pass supplies the laziness the agent lacks.
Use red/green/refactor TDD. Test Driven Development is unpopular with most engineers, except those for whom it’s incredibly popular. I rarely found it useful myself. With agents, that changed. Agents tend to write what Thorsten Ball calls scared mouse code. What if this is null? Missing? Undefined? Every path gets wrapped in another defence until confidence and clarity disappear. You can end up with code that throws no error and does nothing and it’s miserable to debug. Try simply telling the agent “use red/green TDD” and it will write the tests, watch them fail, implement the behaviour, then refactor. That constrains the work and makes success observable. Pro tip: the tests that guided the agent needn’t all survive into the PR. They are merely instrumental. Consolidate redundant micro-tests. Keep high-signal behavioural and E2E tests. TDD is also increasingly reinforced into the agents, so consider the opposite advice and sometimes just tell them to “YOLO with no tests” (especially for prototypes).
Prompting heuristics. Constrain before you delegate. These prompts help me establish the shape of the problem:
- “Interview me until the requirements are clear.”
- “Explain the trade-offs before recommending an approach.”
- “Inspect the repo and explain the current shape before changing anything.”
- “Propose the smallest correct change. Do not edit yet.”
Have them write simply. Reasoning models increasingly speak in convoluted mannered jargon. I suspect this is because RLVR post-training blurs the boundary between user-facing prose and internal chain-of-thought neuralese. They are “RLFried.” Claude Opus 5 and Fable 5/5.1 were particularly bad in my experience. Anthropic acknowledges denser prose in Fable 5.1 and they appear to be actively working on solution. There are some tricks to mitigate this in the near term. My go-to prompt snippet (in most of my system prompts too):
Write all responses, comments, docs, PRs in ultraconcise and ultraclear Simplified Technical English (ASD-STE100). No mannered speech. No filler words or hedging. Lead with the answer. Be direct and clear. Max 15 words per sentence. One idea per sentence. No parenthetical asides. No em-dashes. No semi-colons (unless listing). Every word of text must strongly justify its existence.
Go forth and wield the magic well
Remember, at this point in time coding agents primarily automate the slow but easy parts. The hard and important parts are deciding, verifying, understanding, editing. Those are yours.
That tension is your craft now. And soon it may all change once again.
Godspeed.