Skip to content
Tí Root
Go back

Harness Engineering là gì và vì sao nó là môn engineering quan trọng nhất 2026

Edit page

Một team OpenAI ship một triệu dòng code production mà không viết tay dòng nào. Anthropic publish 3 bài paper. ThoughtWorks formalize framework. Philipp Schmid gọi nó là môn engineering quan trọng nhất 2026. Tên: Harness Engineering.

Table of contents

Open Table of contents

Harness là gì

Harness là yên cương, không phải con ngựa

Agent = Model + Harness. Harness = constraints + feedback + docs + tools. Strip bỏ → model đoán mò. Thêm đúng → ship code production.

Analogy: OS

Model = CPU, Harness = OS

Model = CPU, context = RAM, harness = OS, agent = application. Hầu hết team chạy app không có OS — đó là lý do fail ở production.

2026: môi trường thắng model

Từ prompt engineering sang harness engineering

Cùng model, LangChain đổi harness: 52.8% → 66.5%. Vercel bỏ 80% tools, tốt hơn. The agent was never the hard part — the harness is.

1. AGENT.md / CLAUDE.md

File markdown rải khắp codebase

File markdown rải codebase, agent đọc đầu session. OpenAI AGENT.md, Anthropic CLAUDE.md, Cursor .cursorrules. Một file mỗi module lớn.

2. JSON Feature Lists

Progress tracker trong một file JSON

File JSON track feature + verify + pass/fail. Agent pick feature fail ưu tiên cao nhất, implement, đánh pass, commit. JSON ít bị overwrite ngoài ý muốn hơn Markdown trong run dài.

3. Session Initialization

Boot sequence 7 bước cố định

Boot 7 bước cố định: confirm dir → git log → feature list → start dev → E2E → implement 1 feature → commit. Không có → phí 20 phút đầu session.

4. Sprint Contracts

Hai agent negotiate trước khi code

Generator đề xuất build/pass. Evaluator review. Chỉ khi đồng ý mới implement. Plan + execute cùng pass = output không đáng tin.

5. Structured Task Templates

Grounded impact map trước khi code

Harness analyze codebase thật, sinh impact map với file path + symbol name thật. Tránh hallucinate.

OpenAI: Environment-First

Strict dependency flow 6 tầng

Thiết kế môi trường kỹ đến mức agent tự sinh output review-able: Types → Config → Repo → Service → Runtime → UI, AGENT.md, CI/CD. Sora Android: 4 engineers, 28 ngày, #1 Play Store.

Anthropic: tách Doer khỏi Judge

Ba agent: Planner, Generator, Evaluator

Self-eval không work. Ba agent: Planner (prompt → spec), Generator (1 sprint = 1 feature), Evaluator (browser automation). Solo $9/20 phút broken, harness $200/6 tiếng chạy đúng.

ThoughtWorks: 2×2 Framework

Hai trục: khi nào × cách nào

Mọi control = 2 trục: khi nào chạy (feedforward/feedback) × cách nào chạy (computational/inferential). Cần cả hai, layered.

1. Context đánh bại instructions

Show thực tế, đừng nói abstract

Show trạng thái thực tế luôn outperform nói abstract. Code grounded trong file path thật đánh bại code từ mô tả mơ hồ.

2. Plan và execute phải tách ra

Hard gate giữa planning và implementation

Plan + execute cùng pass = output không đáng tin. Phải là bước riêng, có output review trước khi implement.

3. Feedback loops là bắt buộc

Sensors sau khi agent hành động

Harness không có feedback = prompt thêm vài bước. OpenAI dùng CI, Anthropic LLM khác, ThoughtWorks bảo dùng cả hai.

4. Một thứ tại một thời điểm

Forced incrementalism

Multitask → cạn context, mất coherence, drop requirements. Một feature mỗi sprint, một commit sau mỗi cái.

5. Codebase IS the documentation

Repo là single source of truth

Repo là single source of truth. Convention không có trong codebase = agent không biết.

Harness Decay

Mỗi model upgrade xóa bớt component cũ

Component encode giả định về việc model không làm được gì. Model tốt hơn → giả định hết hạn → overhead. 4.5 cần sprint decomposition, 4.6 bỏ (tiết kiệm 38% cost), 4.7 tự verify.

Build to Delete

Test component bằng cách tắt đi

Thiết kế component để gỡ. Tắt định kỳ, đo, không đổi thì xóa. Manus 5 lần/6 tháng, LangChain 3 lần/năm.

Cost Reality

Solo $9 vs full harness $200

Solo $9/20 phút broken, full harness $200/6 tiếng chạy đúng. 22× cost. Sau một model upgrade: $200 → $124.

Tổng kết

Công thức, 3 trường phái, 5 nguyên lý, build to delete

Kỹ sư thắng 2026 không viết code hay nhất. Họ thiết kế constraints tốt nhất — và vứt đi khi ngừng earning their keep. Bắt đầu từ một AGENT.md. Đo. Mở rộng.


Based on Rahul’s (@sairahul1) June 7, 2026 article on X, this rewrite keeps the 20 source visuals and condenses the prose: the environment an agent runs in is now a bigger engineering problem than the model itself.


Edit page
Share this post on:

Previous Post
Outer Harness — Tại sao Process và Data quan trọng hơn Agent
Next Post
Dynamic Workflows Are a Better Harness for Agentic Work