Harness Engineering: Engineering Reliability for AI Agents

Designing instructions, state and verification around Claude Code, Codex and Cursor — from diagnostics to autonomous graphs

15 bài giảng 5g 57phút Trong đăng ký
Khóa học này dành cho ai
Developers and technical leads who already work with an AI coding agent (Claude Code, OpenAI Codex, Cursor, Gemini CLI, Aider) but get inconsistent results. Ideal for those who want their agent to run productively for hours and across sessions rather than ten minutes under close supervision, and for anyone building internal tooling or a platform around agents.
Yêu cầu
Git, the command line, and hands-on experience with at least one AI coding agent. No machine learning background is required — this course is not about how the model works under the bonnet, but about how to engineer a dependable environment around it.

Chương trình học

15 bài giảng
1
Giới thiệu Your Agent Isn't Broken — What's Around It Is
7 phút
Miễn phí Xem
Why 95% per-step accuracy kills an agent by the tenth step
We lay a formal foundation for harness engineering: the agent as a feedback system (perception → decision → action → evaluation), the mathematics of error accumulation over multi-step runs, Ashby's Law (requisite variety) and Goodhart's Law (a metric that becomes a target ceases to be a metric), context as a noisy channel, and contracts, invariants and postconditions as the only objective criterion of readiness.
Agent control loop Error accumulation across chained steps Ashby's Law Goodhart's Law Context as a noisy channel Contracts and postconditions
26 phút
Sau đăng ký Đăng ký
What Actually Measures Agent Reliability, Rather Than 'Vibes'
We continue building the theoretical foundations: MTBF/MTTR and blast radius as metrics of harness reliability; queues/WIP and Little's Law as a model for why an agent's parallel tasks hold each other up; the Markovian nature of session state (the future depends only on the current state, not on history); confidence calibration (when an agent lies 'confidently'); entropy and Lehman's laws as applied to harness degradation over time; and a consolidation of all this theory into a single table for rapid diagnosis.
MTBF/MTTR blast radius Little's Law WIP (work in progress) Markovian session state confidence calibration harness entropy Lehman's laws
24 phút
Trong đăng ký Đăng ký ngay
The Most Powerful Model on the Market — and It Still Makes Cock-Ups in Production
We unpack why even flagship models routinely fail at seemingly simple tasks in real-world repositories. We introduce the five layers where an agent actually breaks down (task, instructions, state, tools, verification), together with a diagnostic loop that pinpoints the failing layer in a single pass.
the five layers of failure the diagnostic loop model capability versus execution reliability
22 phút
Trong đăng ký Đăng ký ngay
The five parts that make up an agent which actually works
Five harness subsystems — instructions, tools, environment, state and feedback loops — and how they interact. We examine the typical adoption curve (what to fix first), a method for measuring the value of each component via ablation with a fixed model, and the resulting diagram of flows between subsystems.
the five harness subsystems ablation measurement adoption curve
24 phút
Trong đăng ký Đăng ký ngay
Your Agent Is Making Up the Project Structure Because You Never Left It a Map
The repository is the single source of truth for an agent with no memory between sessions. We examine what to measure when mapping your repository, four principles of a good map, a workable file structure, and ACID-like guarantees for agent state.
the repository as the single source of truth the repository map ACID for agent state
24 phút
Trong đăng ký Đăng ký ngay
The rule is written right there in the file. The agent doesn't read it. Why?
Why one giant instructions file is worse than having none at all. Four independent causes of instruction degradation, the SNR (signal-to-noise ratio) metric for instructions, a three-tier architecture of router → documents → executable checks, the rule passport, and practical guidelines for keeping instructions in good working order.
router architecture instruction SNR three-tier architecture rule passport
26 phút
Trong đăng ký Đăng ký ngay
Your Agent Has Amnesia Every Session — Here's the Treatment Plan
An agent is like a brilliant engineer with complete amnesia every morning: why does each new session re-ask questions that were settled yesterday and contradict its own earlier decisions? We examine compaction versus context reset, four continuity artefacts (PROGRESS.md, session handoff, git checkpoints and a decision log), plus the recovery cost metric. We also cover initialisation as a phase in its own right: a bootstrap contract built on four conditions, an acceptance checklist, warm start versus cold start, and a minimal scripts/init.sh.
cross-session continuity session handoff bootstrap contract warm start
28 phút
Trong đăng ký Đăng ký ngay
Three features at once, none finished: the story of one agent
Why the agent has no brakes, and what happens when it grabs hold of three tasks at once — and how a hard WIP=1 limit and clear task boundaries put things right. Then, the feature list (feature_list.) as a genuine harness primitive: the three-part structure of each behavior/verification/state record, a four-state finite state machine, scripts/verify.sh as the only party entitled to set a status to passing, and the difference between strict and weak verification.
WIP=1 task granularity feature_list. feature finite state machine gatekeeper
28 phút
Trong đăng ký Đăng ký ngay
The agent wrote 'Done ✅'. The feature doesn't work. Let's work out why
The verification gap: the divide between 'the agent said it's done' and 'the feature actually works'. Three-layer completion validation, a Definition of Done in AGENTS.md, priority constraints (no refactoring before verification), red flags for the agent — self-healing error messages, and a three-role architecture (implementer / verifier / arbiter).
verification gap Definition of Done three-layer validation separation of roles
24 phút
Trong đăng ký Đăng ký ngay
All Tests Green — Yet the Feature Is Broken. Where Is the Bug Hiding?
Blind spots in unit testing: an agent passes every module test, yet the feature fails to work as a whole. Why E2E tests alter the agent's behaviour before they even catch the bug. How to turn an architectural rule ('the service must not access the database directly') into an executable check rather than a paragraph of text. The verification hierarchy, from linting all the way up to E2E.
Isolation blind spots End-to-end (E2E) tests Executable architectural rules Verification hierarchy
26 phút
Trong đăng ký Đăng ký ngay
The Agent Did the Work — Yet Nobody Can Explain What It Did
Everything appears to work, yet no one — neither a human nor the agent's next session — can explain what exactly was done and why. We break down the two layers of observability, the sprint contract (in scope / out of scope / open questions settled before work begins) and an evaluator rubric in place of general opinion. Then comes clean handover and the battle against entropy: five dimensions of clean state, a session exit checklist, two-mode cleanup and idempotent cleanup operations.
harness observability sprint contract evaluator rubric clean handover harness entropy
28 phút
Trong đăng ký Đăng ký ngay
How to let an agent work for hours without your supervision
A shift from one-off prompting to an autonomous loop. The simplest possible loop; four types of loops that must not be confused; six loop primitives; the separation of generator and evaluator as the only non-negotiable guarantee; an example of a well-designed loop (goal, metric, constraints, order of operations, stopping conditions, escalation); four quiet costs of autonomy; and the maturity ladder.
Loop Engineering generator and evaluator stopping conditions escalation to a human
26 phút
Trong đăng ký Đăng ký ngay
When a Single Loop Is No Longer Enough for an Agent
Where the limits of a single-loop agent lie: three structural failures (no rollbacks, no shared state, no anchors to reality). A four-level framework and four elements of a graph. Anchors — fixed points that cannot be bypassed by local optimisation. The difference between a graph and a workflow. How to build your first graph in six steps. The orchestration tax, and five criteria for when a graph is genuinely needed.
Graph Engineering anchors orchestration tax graph versus workflow
24 phút
Trong đăng ký Đăng ký ngay
Assembling the Harness in Full: From an Empty Repository to an Autonomous Agent
Course finale: we take everything we built step by step across the end-to-end project — a payments FastAPI backend — and bring it together into one working system. We walk through every harness artefact from start to finish: the AGENTS.md router with its topic documents, feature_list. with its finite state machine, PROGRESS.md and the shift-handover protocol, scripts/verify.sh and scripts/arch-check.sh as verification gates, the sprint contract and assessor rubric, the clean session-exit checklist, program.md for the autonomous loop and graph.md for the multi-role process. To finish, we compare the agent's metrics before and after adopting the harness.
assembling the harness as a whole end-to-end project before/after metrics capstone
20 phút
Trong đăng ký Đăng ký ngay