Một team OpenAI ship một triệu dòng code production mà không viết tay dòng nào. Anthropic publish 3 bài paper. ThoughtWorks formalize framework. Philipp Schmid gọi nó là môn engineering quan trọng nhất 2026. Tên: Harness Engineering.
Table of contents
Open Table of contents
- Harness là gì
- Analogy: OS
- 2026: môi trường thắng model
- 1. AGENT.md / CLAUDE.md
- 2. JSON Feature Lists
- 3. Session Initialization
- 4. Sprint Contracts
- 5. Structured Task Templates
- OpenAI: Environment-First
- Anthropic: tách Doer khỏi Judge
- ThoughtWorks: 2×2 Framework
- 1. Context đánh bại instructions
- 2. Plan và execute phải tách ra
- 3. Feedback loops là bắt buộc
- 4. Một thứ tại một thời điểm
- 5. Codebase IS the documentation
- Harness Decay
- Build to Delete
- Cost Reality
- Tổng kết
Harness là gì

Agent = Model + Harness. Harness = constraints + feedback + docs + tools. Strip bỏ → model đoán mò. Thêm đúng → ship code production.
Analogy: OS

Model = CPU, context = RAM, harness = OS, agent = application. Hầu hết team chạy app không có OS — đó là lý do fail ở production.
2026: môi trường thắng model

Cùng model, LangChain đổi harness: 52.8% → 66.5%. Vercel bỏ 80% tools, tốt hơn. The agent was never the hard part — the harness is.
1. AGENT.md / CLAUDE.md

File markdown rải codebase, agent đọc đầu session. OpenAI AGENT.md, Anthropic CLAUDE.md, Cursor .cursorrules. Một file mỗi module lớn.
2. JSON Feature Lists

File JSON track feature + verify + pass/fail. Agent pick feature fail ưu tiên cao nhất, implement, đánh pass, commit. JSON ít bị overwrite ngoài ý muốn hơn Markdown trong run dài.
3. Session Initialization

Boot 7 bước cố định: confirm dir → git log → feature list → start dev → E2E → implement 1 feature → commit. Không có → phí 20 phút đầu session.
4. Sprint Contracts

Generator đề xuất build/pass. Evaluator review. Chỉ khi đồng ý mới implement. Plan + execute cùng pass = output không đáng tin.
5. Structured Task Templates

Harness analyze codebase thật, sinh impact map với file path + symbol name thật. Tránh hallucinate.
OpenAI: Environment-First

Thiết kế môi trường kỹ đến mức agent tự sinh output review-able: Types → Config → Repo → Service → Runtime → UI, AGENT.md, CI/CD. Sora Android: 4 engineers, 28 ngày, #1 Play Store.
Anthropic: tách Doer khỏi Judge

Self-eval không work. Ba agent: Planner (prompt → spec), Generator (1 sprint = 1 feature), Evaluator (browser automation). Solo $9/20 phút broken, harness $200/6 tiếng chạy đúng.
ThoughtWorks: 2×2 Framework

Mọi control = 2 trục: khi nào chạy (feedforward/feedback) × cách nào chạy (computational/inferential). Cần cả hai, layered.
1. Context đánh bại instructions

Show trạng thái thực tế luôn outperform nói abstract. Code grounded trong file path thật đánh bại code từ mô tả mơ hồ.
2. Plan và execute phải tách ra

Plan + execute cùng pass = output không đáng tin. Phải là bước riêng, có output review trước khi implement.
3. Feedback loops là bắt buộc

Harness không có feedback = prompt thêm vài bước. OpenAI dùng CI, Anthropic LLM khác, ThoughtWorks bảo dùng cả hai.
4. Một thứ tại một thời điểm

Multitask → cạn context, mất coherence, drop requirements. Một feature mỗi sprint, một commit sau mỗi cái.
5. Codebase IS the documentation

Repo là single source of truth. Convention không có trong codebase = agent không biết.
Harness Decay

Component encode giả định về việc model không làm được gì. Model tốt hơn → giả định hết hạn → overhead. 4.5 cần sprint decomposition, 4.6 bỏ (tiết kiệm 38% cost), 4.7 tự verify.
Build to Delete

Thiết kế component để gỡ. Tắt định kỳ, đo, không đổi thì xóa. Manus 5 lần/6 tháng, LangChain 3 lần/năm.
Cost Reality

Solo $9/20 phút broken, full harness $200/6 tiếng chạy đúng. 22× cost. Sau một model upgrade: $200 → $124.
Tổng kết

Kỹ sư thắng 2026 không viết code hay nhất. Họ thiết kế constraints tốt nhất — và vứt đi khi ngừng earning their keep. Bắt đầu từ một AGENT.md. Đo. Mở rộng.
Based on Rahul’s (@sairahul1) June 7, 2026 article on X, this rewrite keeps the 20 source visuals and condenses the prose: the environment an agent runs in is now a bigger engineering problem than the model itself.