Research 01 ·
The irreversible costing shift in software engineering
Writing code got cheap. Judgment did not.
Between 2025 and 2026, frontier labs turned software execution into something closer to a bulk commodity. The limited input is human attention spent steering agents, verifying what they produce, and designing the harnesses that make fleets trustworthy.
Multi-page walkthrough of the costing shift, attention reweight, capability, unit economics, cost versus capability, org flip, loops, and harness playbook. Scroll the thumbnails, step back and forth, then open full screen to present.
Slide 01 of 11·Title
Double-click a thumbnail or use Open full screen to present. Arrow keys work in full screen.
For most of software history, delivery cost was dominated by the labor of writing and debugging code. In 2025–2026 that assumption broke. Agents made the mechanical loop (draft, run, fix, repeat) dramatically cheaper and massively parallel. Product sense, verification, and system design stayed expensive in human hours.
Call that the costing shift. It is irreversible for the same reason cloud hardware got cheaper. Capability and unit economics both point the same way. Organizations that still budget engineers primarily as authors of lines will mis-allocate the most limited resource they have, which is people. Builders still matter; the job shifts toward judgment, system design, and the fleets those builders command.
1 · Leaving the coding bottleneck
Limited human hours sit downstream of the agent loop. When agents absorb draft→test→fix work, typing stops setting the schedule. The queue moves to work that still requires judgment, including review, specification, coordination, and coherence.
The bottleneck moved from coding to everything around coding.
Human-first (pre-agent)
Agent-first (high adoption)
- Implementation / coding
Agents absorb draft→test→fix loops; humans keep residual edits and hard edges.[7][8]
Human48%Agent12%Δ-36 pp - Review & verification
More / larger PRs raise review load; verification becomes the long pole.[5]
Human15%Agent34%Δ+19 pp - Specification & design
Ambiguous goals get more expensive when failed attempts are cheap to produce.[1][6]
Human14%Agent22%Δ+8 pp - Coordination & orchestration
Fleet command covers dispatch, triage, stop conditions, and multi-agent handoffs.[1][4]
Human13%Agent18%Δ+5 pp - Coherence, architecture & harness
System integrity, evals, and harness design replace line authorship as leverage.[4]
Human10%Agent14%Δ+4 pp
Human hours stay the limited input. According to Figure 1, implementation falls from 48% to 12% of human attention (−75%), while review and verification rise from 15% to 34% (+127%). Industry telemetry tracks the same pressure: higher PR volume, larger diffs, and longer review cycles under high AI adoption[5][9]. When agents absorb draft and patch work, the schedule is set by judgment capacity rather than typing speed. Agents reweight limited human hours toward the judgment work that always existed. That reallocation holds only if agents also got much more capable and much cheaper to run. If execution stayed limited, teams would hire more authors and the keyboard would remain the long pole.
2 · Raising capability as execution got cheaper
In 2024–2026 capability and cost moved in the same direction. Frontier resolve rates climbed, labs reported order-of-magnitude more agent-authored output, and list and platform token prices fell[18]. The API cost of a strong score still varies widely by model and harness. Stronger models make bulk agent authorship possible. Lower unit costs make that authorship affordable at fleet scale. Capability, volume, and price rewrite the delivery budget only when they reinforce one another.
Humans steer. Agents execute.
Data table and sources
Full row metrics and source notes for the figure above.
- 22%
- 41.3%
- 49%
- 70.3%
- 72.7%
- 74.9%
- 80.9%
- 80%
- 77.2%
- 82.6%
- 88.6%
- 78.2%
- 86.6%
- 93.4%
- 96.2%
| Quarter | Model | Resolve % | Source |
|---|---|---|---|
| 2024 Q1 | Claude 3 Opus | 22% | vendor Verified · Mar 2024[11] |
| 2024 Q3 | o1-preview | 41.3% | vendor Verified · Sep 2024[12] |
| 2024 Q4 | Claude 3.5 Sonnet | 49% | vendor Verified · Oct 2024 model[11] |
| 2025 Q1 | Claude 3.7 Sonnet | 70.3% | vendor Verified · Feb 2025[13] |
| 2025 Q2 | Claude Sonnet 4 | 72.7% | vendor Verified · May 2025[14] |
| 2025 Q3 | GPT-5 | 74.9% | vendor Verified · Aug 2025[15] |
| 2025 Q4 | Claude Opus 4.5 | 80.9% | vendor Verified · 2025[16] |
| 2026 Q1 | GPT-5.2 Thinking | 80% | vendor Verified · Dec 2025 / Q1 2026[36] |
| 2026 Q1 | GPT-5.4 | 77.2% | Vals Mini-SWE · Mar 2026[37] |
| 2026 Q2 | GPT-5.5 | 82.6% | Vals Mini-SWE · Apr 2026[38] |
| 2026 Q2 | Claude Opus 4.8 | 88.6% | Vals Mini-SWE · May 2026[10] |
| 2026 Q2 | Kimi K2.7 Code | 78.2% | Vals Mini-SWE · Jun 2026[41] |
| 2026 Q3 | Grok 4.5 | 86.6% | Vals Mini-SWE · Jul 2026[10] |
| 2026 Q3 | Kimi K3 | 93.4% | Vals Mini-SWE · Jul 2026[42] |
| 2026 Q3 | GPT-5.6 Sol | 96.2% | Vals Mini-SWE · Jul 2026[10] |
According to Figure 2, the frontier envelope climbs from 22% resolve rate in early 2024 to 96.2% by mid-2026 (+337%), then packs several models into the same high band near the ceiling. The line is the highest resolve rate at each date. Lower dots show capability became a pack race across labs.
That pattern supports the costing thesis. Agent execution quality is abundant relative to typing-era limits. When the frontier keeps climbing and alternatives stack on the same calendar, organizations that still budget human attention as if coding were the long pole will underfund review, specification, and harness work, the limited inputs that now set the schedule.
Capability keeps climbing while human review capacity does not scale the same way. Not every release is a new peak, and vendor harnesses differ from public Mini-SWE runs, yet agent execution is more than a temporary spike. Independent time-horizon measures show agent task length roughly doubling on a short cycle, the same direction as the resolve-rate ladder[21].
According to Figure 3, the four tiles move as one direction of travel. Calendar time falls by about 10× in the OpenAI harness report, Claude-authored merges reach 80%+, merge volume per engineer is about 8× higher than 2024, and a token-cost sample drops about −67% year over year. Treat the multipliers as quantity proxies from labs dogfooding their own tools, not a universal productivity KPI. The joint pattern still matters.
Together with the resolve-rate ladder, that pattern is why reallocation of human attention sticks. When retries and bulk authorship get cheaper, process waste shows up as limited human hours on review, intent, and orchestration rather than calendar typing time. Suite cost against resolve rate then decides whether the leverage stays affordable at fleet scale.
Volume and token cost tell half the economics story. The other half is whether a strong score is cheap enough to run as a default fleet path with model and harness together, not resolve rate alone. That cost–capability surface is the real fleet decision.
Benchmark Type:
Filter:
Data table and sources
Full row metrics and source notes for the figure above.
- GPT-5.6 Sol · max[10]
max· Mini-SWE-agent · SWE-bench VerifiedResolve96.2%$/task$1.15Suite $$575List $5/$30 in/out per MTok - Fable 5 · max[10]
max· Mini-SWE-agent · SWE-bench VerifiedResolve95%$/task$2.05Suite $$1025 - Kimi K3[10]
default· Mini-SWE-agent · SWE-bench VerifiedResolve93.4%$/task—Suite $—List $3/$15 in/out per MTok - GPT-5.6 Luna · max[10]
max· Mini-SWE-agent · SWE-bench VerifiedResolve93%$/task$0.21Suite $$105List $1/$6 in/out per MTok - Opus 4.8 · max[10]
max· Mini-SWE-agent · SWE-bench VerifiedResolve88.6%$/task$1.92Suite $$960List $5/$25 in/out per MTok - Grok 4.5 · high[10]
high· Mini-SWE-agent · SWE-bench VerifiedResolve86.6%$/task$0.54Suite $$270List $2/$6 in/out per MTok - Opus 4.8 · Claude Code[10]
default· Claude Code · SWE-bench VerifiedResolve85.8%$/task$0.67Suite $$335List $5/$25 in/out per MTok - GLM 5.2 · default[10]
default· Mini-SWE-agent · SWE-bench VerifiedResolve82.8%$/task$0.71Suite $$355 - Sonnet 5 · max[10]
max· Mini-SWE-agent · SWE-bench VerifiedResolve79.6%$/task$1.49Suite $$745List $3/$15 in/out per MTok - Kimi K2.7 Code[10]
default· Mini-SWE-agent · SWE-bench VerifiedResolve78.2%$/task$0.27Suite $$135List $0.95/$4 in/out per MTok - Sonnet 4.6 · max[10]
max· Mini-SWE-agent · SWE-bench VerifiedResolve77.4%$/task$1.30Suite $$650List $3/$15 in/out per MTok - Opus 4.5 · high[10]
high· Mini-SWE-agent · SWE-bench VerifiedResolve76.8%$/task$0.75Suite $$375List $5/$25 in/out per MTok - GPT-5.6 Terra · xhigh[10]
xhigh· Mini-SWE-agent · SWE-bench VerifiedResolve75.2%$/task$0.72Suite $$360List $2.5/$15 in/out per MTok - Sonnet 4.5 · high[10]
high· Mini-SWE-agent · SWE-bench VerifiedResolve71.4%$/task$0.66Suite $$330List $3/$15 in/out per MTok - Haiku 4.5 · high[10]
high· Mini-SWE-agent · SWE-bench VerifiedResolve66.6%$/task$0.33Suite $$165
| Model | Effort | Harness | Resolve % | $/task | Suite total ($ · 500) | List price in/out |
|---|---|---|---|---|---|---|
| GPT-5.6 Sol · max[10] | max | Mini-SWE-agent | 96.2% | $1.15 | $575 | $5/$30 |
| Fable 5 · max[10] | max | Mini-SWE-agent | 95% | $2.05 | $1025 | — |
| Kimi K3[10] | default | Mini-SWE-agent | 93.4% | — | — | $3/$15 |
| GPT-5.6 Luna · max[10] | max | Mini-SWE-agent | 93% | $0.21 | $105 | $1/$6 |
| Opus 4.8 · max[10] | max | Mini-SWE-agent | 88.6% | $1.92 | $960 | $5/$25 |
| Grok 4.5 · high[10] | high | Mini-SWE-agent | 86.6% | $0.54 | $270 | $2/$6 |
| Opus 4.8 · Claude Code[10] | default | Claude Code | 85.8% | $0.67 | $335 | $5/$25 |
| GLM 5.2 · default[10] | default | Mini-SWE-agent | 82.8% | $0.71 | $355 | — |
| Sonnet 5 · max[10] | max | Mini-SWE-agent | 79.6% | $1.49 | $745 | $3/$15 |
| Kimi K2.7 Code[10] | default | Mini-SWE-agent | 78.2% | $0.27 | $135 | $0.95/$4 |
| Sonnet 4.6 · max[10] | max | Mini-SWE-agent | 77.4% | $1.30 | $650 | $3/$15 |
| Opus 4.5 · high[10] | high | Mini-SWE-agent | 76.8% | $0.75 | $375 | $5/$25 |
| GPT-5.6 Terra · xhigh[10] | xhigh | Mini-SWE-agent | 75.2% | $0.72 | $360 | $2.5/$15 |
| Sonnet 4.5 · high[10] | high | Mini-SWE-agent | 71.4% | $0.66 | $330 | $3/$15 |
| Haiku 4.5 · high[10] | high | Mini-SWE-agent | 66.6% | $0.33 | $165 | — |
According to Figure 4, the scatter is a fleet decision surface. Up is better and left is cheaper. On Verified Mini-SWE, systems can sit near 93–96% resolve rate while unit economics diverge: GPT-5.6 Luna at $0.21/task is about −82% cheaper per task than GPT-5.6 Sol at $1.15/task for 93.0% vs 96.2% resolve. DeepSWE pulls the cloud down and spreads dollars further. Harder work separates a strong score from a package cheap enough to run at scale.
That spread is industry evidence for the costing shift. Capability without unit economics is incomplete. The hard judgment is which model–effort–harness package stays leftward for default automation. Teams that chase headline resolve % alone overpay on volume loops and underfund the human review and design work those loops create, and that reallocation of human hours is the irreversible move.
The frontier climbed and lab volume exploded, but model and harness selection still determines whether that leverage is economically sustainable[10][43]. Once that economics is real, org design has to change. Leave the old headcount model in place and the bottleneck simply moves into review queues and coordination debt.
3 · Transforming to AI-native organization
When execution is bulk and judgment is limited, headcount plans and rituals built for line authorship miss the queue. An AI-native engineering org flips the default. Agents build. Humans keep the work that compresses poorly, such as product sense, architecture legibility, eval design, taste, and the harnesses that make fleets trustworthy.
Coding is no longer the bottleneck.
According to Figure 5, the limited input flips. When typing is limited, labor owns the schedule, code is the long pole, heavy planning feels rational, and review tracks human authorship. When execution is bulk, judgment is limited, verification sets the calendar, evals and harnesses buy leverage, and review tracks agent throughput unless the org redesigns it. The operating model follows that economics. AI-native teams treat agents as default builders and keep humans on intent, verification, and harness design. “Code wins” over ceremony, overnight agents come with harder review, and smaller cores own more surface[1][2][4]. When overnight agents open a flood of pull requests, the review queue sets the calendar, not typing time. Teams that fold failures into harness rules keep the leverage; teams that only add more authors re-create the old bottleneck in a new shape.
The pattern is leaving the labs. Anthropic × NEC put Claude Code in front of tens of thousands of employees[23]. New roles such as loop designers, eval owners, and fleet operators matter more than planning theater[22]. Those roles earn their keep by designing how work is discovered, dispatched, and verified at scale.
4 · Engineering the loop
When typing is bulk, the high-leverage move is designing systems that prompt, verify, and repeat without a human in every turn.
Loop engineering is replacing yourself as the person who prompts the agent. You design the system that does it instead.
A loop is a recursive, goal-directed system that discovers work, dispatches agents, verifies, records state, and repeats. Mid-2026 made that shift concrete from three directions at once[24]. Boris Cherny, creator of Claude Code, frames the job as writing loops that prompt Claude rather than prompting turn-by-turn[24][25]. Peter Steinberger, creator of OpenClaw, states the same reframe for high-throughput OSS as designing loops that prompt your agents[26]. Andrew Ng, founder of DeepLearning.AI and co-founder of Coursera, lifts loop engineering to product cadence with three development loops at minutes, hours, and days[27][28][29]. The skill moves from prompt craft to system design. Ng’s three product-development loops put those cadences on one map.
According to Figure 6, product work runs on three nested cadences. Agentic coding sits on the order of minutes, developer feedback on the order of hours, and external feedback on the order of days, with product specs and developer vision as shared nodes. Those cadences need runtime primitives before they become a daily loop. Across Codex-style apps and Claude Code the shared stack includes scheduled automations and goals, git worktrees for parallel isolation, Skills (SKILL.md) as progressive project knowledge, MCP/connectors for real tools, sub-agents so maker ≠ checker, and durable state on disk (markdown, tickets, AGENTS.md) because models forget[24]. A typical daily loop runs as a short pipeline.
According to Figure 7, a daily loop is Triage → Isolate → Draft → Verify → Ship, with Human on exceptions and merge only. That pattern stays safe only when judgment stays in the loop. Loops amplify judgment, good or bad. Stop reading output and you accumulate comprehension debt[30]. Verification still ends with a human who owns the merge. And a loop is only as good as the environment it runs in, including docs, tools, tests, hooks, and memory that agents can actually use.
5 · Engineering the harness
A loop without a harness is chat with extra steps. What people experience as “the agent” is the model plus everything wrapped around it. Agent = Model + Harness.
Agent = Model + Harness
System prompts,
CLAUDE.md,AGENTS.md, skill files, and subagent promptsTools, skills, MCP servers, and their descriptions
Bundled infrastructure such as filesystem, sandbox, and browser
Orchestration logic such as subagent spawning, handoffs, and model routing
Hooks and middleware for deterministic execution, including compaction, continuation, and lint checks
Observability such as logs, traces, and cost and latency metering
According to Figure 8, the coding agent is only half the system. The human steers guides and sensors; feedforward and feedback close the loop into the agent; durable state sits outside the context window so failures can ratchet into permanent harness rules. Agent = Model + Harness is the architecture, not a slogan.
OpenAI’s experiment treats the repository knowledge as system of record. A short AGENTS.md is the table of contents (roughly 100 lines), not the encyclopedia. Structured docs/ holds design docs, execution plans, product specs, references, and quality grades. Progressive disclosure keeps the map in context and the depth on disk. Per-worktree bootable apps and strict linters whose errors include remediation instructions make the agent runnable[4].
Multi-hour agent runs need the same structure at longer horizons, with clear stop conditions, durable state, and feedback that keeps unattended work on track[32]. Conference practice adds hooks (red squigglies for agents) and a ratchet mindset. Every observed failure becomes a permanent harness rule[2][31].
AGENTS.md is the table of contents. Structured docs/ is the system of record. Agents open only the branch they need.[4]According to Figure 9, short AGENTS.md is the map injected into context, and structured docs/ is the encyclopedia on disk. Progressive disclosure means the agent opens only the branch it needs. That layout also decides what the model re-reads on every turn. System prompts, tool descriptions, skill files, and long agent maps land in the context window before the task does. Dump the full wrap on every step and tokens burn, useful task context gets crowded out, and the prompt prefix changes enough that platforms cannot reuse work. A production harness designs that load on purpose.
Short
AGENTS.mdmaps into progressivedocs/. The agent starts with a table of contents and opens depth only when the task needs it.Keep system prompts and tool schemas byte-identical across turns. A stable prefix is what platforms can cache; churn kills hit rates.
Shrink live history with compaction and continuation hooks. The window holds the working set. Repository knowledge stays on disk as the durable system of record.
Keep unused tool schemas and raw tool dumps out of the window. Tool definitions and bulky outputs crowd out the task long before they help.
Inspect tool results programmatically and re-enter only needed fields. One conference example cut token use about 66% by moving tool results from JSON to markdown and dropping unused fields.
Treat prompts, tools, hooks, and metering as the affordability levers. They keep multi-step agent loops viable when retries are the normal path.
Context efficiency is harness engineering under a cost meter. At Code with Claude 2026, Brad Abrams (Claude Platform) put a production floor on that meter. Aim for at least an 80% prompt-cache hit rate before other agent optimizations. Mature product harnesses already clear a higher bar. Cursor, Replit, Perplexity, and Claude Code itself run cache rates in the 90s[2]. Cached tokens are cheaper and faster, and they do not count against rate limits. High hits mean the wrap is stable and lean enough that strong models and frequent agent retries can stay the default production path.
6 · Practicing when execution is cheap
Day to day, the costing shift shows up as operating habits. Invest people and budget in the steps that still set the schedule. The durable investments are loops that keep agents moving without a human on every turn, harnesses that make those loops safe and reusable, and verification capacity that scales with agent throughput. More authors alone re-create the old bottleneck. The habits below turn that investment stance into workflow, harness, and org practice.
Workflow
Route recurring work through agents first. Triage, docs, patches, and summaries are good defaults. Keep humans on the exceptions and the merge decision[1].
Run overnight goals only when tests, hooks, and stop conditions are trustworthy. They must catch bad loops before morning review[2][32].
Prefer isolated git worktrees for parallel agents. One run must not corrupt another working tree or block a human hotfix[4][24].
Budget real calendar time to understand code nobody on the team typed. Comprehension debt compounds when review is only a skim[30].
Harness & loops
Keep agent maps short. Use a table of contents plus pointers, and push detail into progressive
docs/so context stays cheap and current[4].Prefer hooks and computational feedback that include remediation text. Giant static prompts decay; machine checks that explain how to fix stay useful[4][31].
Split maker and checker roles, and define “done” before the loop starts. Unattended runs need a clear exit[24].
Persist tickets, plans, and outcomes outside the context window. Models forget; the repository should not[4][24].
Org & cost
Resolve technical debates with working prototypes. “Code wins” is faster than process theater when agents make experiments cheap[1].
Retire ceremonies that only optimized human typing. Grow verification capacity as agent throughput rises[1][3].
Measure human review time and merge success rate, not only lines merged. Quantity proxies hide judgment work[5][9].
Track token spend and suite cost at fleet scale. “Cheap” retries must not become invisible waste[2][18].
Adopt the habits with the same caution used for the evidence itself. Conference practice, harness write-ups, and loop engineering notes converge on the same investment stance; they do not guarantee every tip fits every codebase.
Code is cheap, show me the Harness and Loop
Holding open questions
Four questions still set the research and operating agenda.
1. Outside the frontier lab
How much of this shift survives legacy codebases, regulated domains, and teams without internal model leverage?
2. Comprehension debt
What metrics capture long-term maintainability when most of the codebase was never held in a human working memory?
3. Harness accounting
How should organizations budget evals, harness engineering, and review capacity relative to token spend?
4. Role redesign
What durable career paths look like when the center of gravity moves from authorship to fleet command and system design?
References
- Fiona Fung (). “Running an AI-native engineering org”. Code with Claude 2026 (YouTube). “The bottleneck moved from coding to everything around coding”; Claudify work; code wins over process.
- Chris Ebert (). “Notes from Code with Claude 2026”. chrisebert.net. Near-primary conference notes: Fung bottleneck line; Abrams ≥80% prompt-cache floor; Cursor/Replit/Perplexity/Claude Code cache rates in the 90s; hooks as red squigglies; overnight agents.
- Fiona Fung (). “Interview on Lenny’s Podcast”. Lenny’s Podcast / YouTube. “Coding is no longer the bottleneck” organizational diagnosis.
- Ryan Lopopolo / OpenAI (). “Harness engineering: leveraging Codex in an agent-first world”. OpenAI. Internal ~1M LOC product with effectively no hand-written code; AGENTS.md as ~100-line table of contents; structured docs/ progressive disclosure; humans steer, agents execute.
- Addy Osmani, director of Google Cloud AI (). “The 80% problem in agentic coding”. addyo.substack.com. Review becomes the long pole as PR volume and size rise with AI adoption.
- Anthropic (). “2026 Agentic Coding Trends Report”. Anthropic resources. Industry shift from writing to reviewing/directing under agent adoption.
- Sonar (). “How much time do developers spend actually writing code?”. sonarsource.com. Historical coding-time share (~32%) before the agent reweight.
- Software.com / Antenna (). “Global code-time sample”. antenna.dev. Independent sample of active coding time vs total developer week.
- Leadership Garden (Faros / DORA-style synthesis) (). “AI: the work moved”. leadership.garden. Synthesis of telemetry: more PRs, longer review, larger diffs under high AI adoption.
- Vals AI (). “SWE-bench Verified leaderboard”. vals.ai. Common Mini-SWE harness scores and Cost/Test for cross-model comparison.
- Anthropic Engineering (). “Raising the bar on SWE-bench Verified with Claude 3.5 Sonnet”. anthropic.com. Upgraded Claude 3.5 Sonnet 49% Verified (Oct 2024 model); comparison table lists Claude 3 Opus 22% and prior SOTA ~45%.
- OpenAI (). “OpenAI o1 System Card”. openai.com. o1-preview SWE-bench Verified 41.3% (Agentless scaffold).
- Anthropic (). “Claude 3.7 Sonnet and Claude Code”. anthropic.com. Claude 3.7 Sonnet vendor SWE-bench Verified 70.3% (high-compute scaffold); 63.7% with a simpler agent.
- Anthropic (). “Introducing Claude 4”. anthropic.com. Claude Sonnet 4 SWE-bench Verified 72.7% at launch.
- OpenAI (). “Introducing GPT-5 for developers”. openai.com. GPT-5 SWE-bench Verified 74.9%; API $1.25/$10 per MTok.
- Anthropic (). “Introducing Claude Opus 4.5”. anthropic.com. Vendor SWE-bench Verified 80.9%; $5/$25 per MTok pricing.
- OpenAI (). “Why we no longer evaluate SWE-bench Verified”. openai.com. Context on Verified as a shared yardstick and limits of fair comparison.
- OpenAI (). “API pricing”. platform.openai.com. List API prices for OpenAI models used in cost comparisons.
- Anthropic Institute (). “When AI builds itself”. Anthropic. 80%+ Claude-authored merges; ~8× merge volume/engineer vs 2024; flags that LOC multipliers overstate true productivity.
- Open Source For You (AI.cc / AICC report recap) (). “Enterprise AI costs crash 67% as open source models and multi-model routing go mainstream”. Open Source For You. Platform sample: blended enterprise token cost $18.40 → $6.07 / MTok (Q1 2025 → Q1 2026).
- METR (). “Time horizons”. metr.org. Agent task time horizons roughly doubling every few months (cited by Anthropic).
- Martin Fowler & Birgitta Böckeler (). “Harness engineering for coding agent users”. martinfowler.com. Feedforward guides vs feedback sensors; harness as leverage when agents write code.
- Anthropic (). “Anthropic and NEC”. anthropic.com. Claude / Claude Code for ~30,000 NEC Group employees; large AI-native engineering push.
- Addy Osmani, director of Google Cloud AI (). “Loop engineering”. addyosmani.com. Defines loop engineering: design systems that prompt agents; primitives table (worktrees, skills, sub-agents).
- Boris Cherny (via circulated clip) (). “On writing loops that prompt Claude”. X / social synthesis. Creator/head of Claude Code: job is to write loops that prompt Claude, not turn-by-turn prompts.
- Peter Steinberger (@steipete) (). “Design loops that prompt your agents”. X. OpenClaw creator: stop prompting coding agents; design loops that prompt them instead.
- Andrew Ng (@AndrewYNg) (). “Loop engineering letter with three-loops diagram”. X. Public post of the three product-development loops diagram used in Figure 6.
- Andrew Ng (). “3 Key Loops for Building 0-to-1 Products with AI Agents”. LinkedIn / The Batch letter. Agentic coding, developer feedback, and external feedback loops at minutes / hours / days scales.
- Andrew Ng (). “Three loops explainer”. YouTube. Video walkthrough of the three nested product-development loops.
- Addy Osmani, director of Google Cloud AI (). “Comprehension debt”. addyosmani.com. Risk of shipping code nobody on the team understands when agents write at scale.
- Addy Osmani, director of Google Cloud AI (). “Agent harness engineering”. addyosmani.com. Ratchet mindset: turn every failure into a permanent harness rule.
- Anthropic Engineering (). “Effective harnesses for long-running agents”. anthropic.com. Harness design for multi-hour agent runs: structure, feedback, stop conditions.
- Anthropic (). “Introducing Claude Sonnet 4.6”. anthropic.com. Sonnet 4.6 coding upgrade; standard $3/$15 pricing.
- Anthropic (). “Introducing Claude Sonnet 5”. anthropic.com. Agentic Sonnet tier; cost-performance curves at effort levels.
- Anthropic (). “Introducing Claude Opus 4.8”. anthropic.com. Opus 4.8 frontier coding model; Vals Mini-SWE 88.6% class results.
- Moonshot AI (). “Kimi K2.7 Code”. kimi.com. Open coding-focused agentic model; Vals Mini-SWE Verified 78.2% · ~$0.27/test.
- Moonshot AI (). “Kimi K3: Open Frontier Intelligence”. kimi.com. 2.8T-class flagship; Vals Mini-SWE Verified 93.4% (Cost/Test not yet published).
- Datacurve (). “DeepSWE leaderboard”. deepswe.datacurve.ai. Contamination-free long-horizon SE benchmark (113 tasks); mini-swe-agent harness; pass@1 and avg cost.