Research 01 ·
The irreversible costing shift in software engineering
Writing code costs less. Judgment still takes time.
In 2025–2026, frontier AI labs made drafting, running, and repairing code faster and less expensive, with many tasks running in parallel. Human attention increasingly determines delivery speed through agent direction, output verification, and harness design.
Slide 01 of 12·Title
Open a slide in full screen to present. Use the arrow keys to move between slides.
Writing and debugging code once accounted for much of the cost of software delivery. Team size, time at the keyboard, and the cost of revisions set the pace. In 2025–2026, agents changed that balance by making the draft, run, and repair loop less expensive and enabling many tasks to run in parallel. Product judgment, verification, and system design continued to demand human time.
This is the shift in software engineering costs. Like the decline in cloud infrastructure costs, it reflects advances in both capability and unit economics that extend beyond any single product roadmap. Teams that budget engineering primarily by code output risk misallocating their most limited resource, human attention. Engineers remain essential, with more of their work devoted to judgment, system design, and coordinating agents.
Agents collapsed the marginal cost of building. They raised the relative cost of deciding what to build, whether it is correct, and whether the system still coheres.
1 · Human attention moves to review
As agents take on drafting, testing, and repair, human work concentrates on review, specification, coordination, and keeping the system coherent. These tasks increasingly determine delivery time. The first sign of the cost shift is therefore a change in how teams with high agent adoption spend their attention.
The bottleneck moved from coding to everything around coding.
Human-first (pre-agent)
Agent-first (high adoption)
- Implementation / codingHuman-first %48%Agent-first %12%Δ pp-36 pp
- Review & verificationHigher volume and larger pull requests raise review load; verification becomes the critical path.[5]Human-first %15%Agent-first %34%Δ pp+19 pp
- Specification & designHuman-first %14%Agent-first %22%Δ pp+8 pp
- Coordination & orchestrationHuman-first %13%Agent-first %18%Δ pp+5 pp
- Coherence, architecture & harnessSystem integrity, evaluation suites, and harness design replace line authorship as leverage.[4]Human-first %10%Agent-first %14%Δ pp+4 pp
According to Figure 1, Implementation falls from 48% to 12% of human attention (−75%), while Review and Verification rises from 15% to 34% (+127%). In this illustrative allocation, most of the attention released from coding moves to review, with further increases in specification, coordination, and system coherence.
Industry telemetry shows similar pressure in teams with high AI adoption, including more pull requests, larger diffs, and longer review cycles.[5][9] As agents draft and repair code, delivery depends more on the team’s ability to evaluate it. Agents concentrate limited human time on judgment.
This change can persist when agents become both more capable and less expensive to run. Together, those improvements make it practical to shift human attention from producing code to directing and verifying it.
2 · Greater capability at lower cost
Public resolve-rate results show agents completing more tasks, while cost samples show retries becoming less expensive. These two trends help explain why the attention shift in Figure 1 can persist.
Humans steer. Agents execute.
Data table and sources
Detailed values and source notes for this figure.
- 22%
- 41.3%
- 49%
- 70.3%
- 72.7%
- 74.9%
- 80.9%
- 80%
- 77.2%
- 82.6%
- 88.6%
- 78.2%
- 86.6%
- 93.4%
- 96.2%
- 97%
- 88.8%
| Quarter | Model | Resolve % | Source |
|---|---|---|---|
| 2024 Q1 | Claude 3 Opus | 22% | Vendor-reported SWE-bench Verified · March 2024[11] |
| 2024 Q3 | o1-preview | 41.3% | Vendor-reported SWE-bench Verified · September 2024[12] |
| 2024 Q4 | Claude 3.5 Sonnet | 49% | Vendor-reported SWE-bench Verified · October 2024 model[11] |
| 2025 Q1 | Claude 3.7 Sonnet | 70.3% | Vendor-reported SWE-bench Verified · February 2025[13] |
| 2025 Q2 | Claude Sonnet 4 | 72.7% | Vendor-reported SWE-bench Verified · May 2025[14] |
| 2025 Q3 | GPT-5 | 74.9% | Vendor-reported SWE-bench Verified · August 2025[15] |
| 2025 Q4 | Claude Opus 4.5 | 80.9% | Vendor-reported SWE-bench Verified · 2025[16] |
| 2026 Q1 | GPT-5.2 Thinking | 80% | Vendor-reported SWE-bench Verified · December 2025 / Q1 2026[36] |
| 2026 Q1 | GPT-5.4 | 77.2% | Vals Mini-SWE leaderboard · March 2026[37] |
| 2026 Q2 | GPT-5.5 | 82.6% | Vals Mini-SWE leaderboard · April 2026[38] |
| 2026 Q2 | Claude Opus 4.8 | 88.6% | Vals Mini-SWE leaderboard · May 2026[10] |
| 2026 Q2 | Kimi K2.7 Code | 78.2% | Vals Mini-SWE leaderboard · June 2026[41] |
| 2026 Q3 | Grok 4.5 | 86.6% | Vals Mini-SWE leaderboard · July 2026[10] |
| 2026 Q3 | Kimi K3 | 93.4% | Vals Mini-SWE leaderboard · July 2026[42] |
| 2026 Q3 | GPT-5.6 Sol | 96.2% | Vals Mini-SWE leaderboard · July 2026 (behind Claude Opus 5 peak)[10] |
| 2026 Q3 | Claude Opus 5 | 97% | Vals Mini-SWE leaderboard · July 2026; Anthropic list price $5/$25 MTok[47] |
| 2026 Q3 | DeepSeek V4 Flash 0731 | 88.8% | Vals Mini-SWE · July 31 2026 release; HF model card DeepSWE 54.4%[48] |
According to Figure 2, the highest reported resolve rate rises from a 22% resolve rate in early 2024 to 97.0% by mid-2026 (+341%), with Claude Opus 5 and GPT-5.6 Sol occupying the high band near the ceiling while open-weight systems such as DeepSeek V4 Flash 0731 sit in the high 80s. The solid line marks the highest reported rate at each date. The lower dots show that capability became a competitive field across laboratories rather than a single permanent lead. Individual releases vary, and differences between vendor harnesses and public Mini-SWE runs limit direct comparison. Independent time-horizon measures show agent task length roughly doubling on a short cycle,[21] which is the same direction of travel as the rise in resolve rates.
Human review capacity is growing more slowly than agent capability. Teams that continue to treat coding as the main constraint risk underfunding review, specification, and harness design. Cost also matters. A high benchmark score becomes useful at scale only when teams can afford to run the system repeatedly.
According to Figure 3, the four indicators point toward faster, less expensive execution. The OpenAI harness report describes builds completing about 10× faster. Claude-authored merges reach 80%+, merge volume per engineer is about 8× the 2024 level, and the token-cost sample falls about −67% year over year.
These are quantity measures from laboratories using their own tools. They can overstate improvements in judgment and should be interpreted as evidence of cheaper retries and greater output, with limited implications for overall productivity.
As code generation and retries become less expensive, delays increasingly arise in review, clarification, and coordination. The economic question is which model and harness can deliver the required capability at a sustainable cost. Teams need that comparison to choose a default system for recurring work.
Benchmark Type:
Filter:
Data table and sources
Detailed values and source notes for this figure.
- Opus 5 · max[10]
max· Mini-SWE-agent · SWE-bench VerifiedScore %97%$/task$1.29Suite $—List $5/$25 in/out per MTok - GPT-5.6 Sol · max[10]
max· Mini-SWE-agent · SWE-bench VerifiedScore %96.2%$/task$1.15Suite $—List $5/$30 in/out per MTok - Kimi K3[10]
default· Mini-SWE-agent · SWE-bench VerifiedScore %93.4%$/task—Suite $—List $3/$15 in/out per MTok - GPT-5.6 Luna · max[10]
max· Mini-SWE-agent · SWE-bench VerifiedScore %93%$/task$0.04Suite $—List $0.2/$1.2 in/out per MTok - Fable 5 · max[10]
max· Mini-SWE-agent · SWE-bench VerifiedScore %95%$/task$2.05Suite $— - GPT-5.6 Terra · xhigh[10]
xhigh· Mini-SWE-agent · SWE-bench VerifiedScore %75.2%$/task$0.29Suite $—List $2/$12 in/out per MTok - Kimi K2.7 Code[10]
default· Mini-SWE-agent · SWE-bench VerifiedScore %78.2%$/task$0.27Suite $—List $0.95/$4 in/out per MTok - Opus 4.8 · max[10]
max· Mini-SWE-agent · SWE-bench VerifiedScore %88.6%$/task$1.92Suite $—List $5/$25 in/out per MTok - Grok 4.5 · high[10]
high· Mini-SWE-agent · SWE-bench VerifiedScore %86.6%$/task$0.54Suite $—List $2/$6 in/out per MTok - Opus 4.8 · Claude Code[10]
default· Claude Code · SWE-bench VerifiedScore %85.8%$/task$0.67Suite $—List $5/$25 in/out per MTok - GLM 5.2 · default[10]
default· Mini-SWE-agent · SWE-bench VerifiedScore %82.8%$/task$0.71Suite $— - Sonnet 5 · max[10]
max· Mini-SWE-agent · SWE-bench VerifiedScore %79.6%$/task$1.49Suite $—List $3/$15 in/out per MTok - Sonnet 4.6 · max[10]
max· Mini-SWE-agent · SWE-bench VerifiedScore %77.4%$/task$1.30Suite $—List $3/$15 in/out per MTok - Opus 4.5 · high[10]
high· Mini-SWE-agent · SWE-bench VerifiedScore %76.8%$/task$0.75Suite $—List $5/$25 in/out per MTok - Sonnet 4.5 · high[10]
high· Mini-SWE-agent · SWE-bench VerifiedScore %71.4%$/task$0.66Suite $—List $3/$15 in/out per MTok - Haiku 4.5 · high[10]
high· Mini-SWE-agent · SWE-bench VerifiedScore %66.6%$/task$0.33Suite $— - DeepSeek V4 Flash 0731[10]
max· Mini-SWE-agent · SWE-bench VerifiedScore %88.8%$/task$0.01Suite $—
| Model | Effort | Harness | Score % | $/task | Suite total ($ · 500) | List price in/out |
|---|---|---|---|---|---|---|
| Opus 5 · max[10] | max | Mini-SWE-agent | 97% | $1.29 | — | $5/$25 |
| GPT-5.6 Sol · max[10] | max | Mini-SWE-agent | 96.2% | $1.15 | — | $5/$30 |
| Kimi K3[10] | default | Mini-SWE-agent | 93.4% | — | — | $3/$15 |
| GPT-5.6 Luna · max[10] | max | Mini-SWE-agent | 93% | $0.04 | — | $0.2/$1.2 |
| Fable 5 · max[10] | max | Mini-SWE-agent | 95% | $2.05 | — | — |
| GPT-5.6 Terra · xhigh[10] | xhigh | Mini-SWE-agent | 75.2% | $0.29 | — | $2/$12 |
| Kimi K2.7 Code[10] | default | Mini-SWE-agent | 78.2% | $0.27 | — | $0.95/$4 |
| Opus 4.8 · max[10] | max | Mini-SWE-agent | 88.6% | $1.92 | — | $5/$25 |
| Grok 4.5 · high[10] | high | Mini-SWE-agent | 86.6% | $0.54 | — | $2/$6 |
| Opus 4.8 · Claude Code[10] | default | Claude Code | 85.8% | $0.67 | — | $5/$25 |
| GLM 5.2 · default[10] | default | Mini-SWE-agent | 82.8% | $0.71 | — | — |
| Sonnet 5 · max[10] | max | Mini-SWE-agent | 79.6% | $1.49 | — | $3/$15 |
| Sonnet 4.6 · max[10] | max | Mini-SWE-agent | 77.4% | $1.30 | — | $3/$15 |
| Opus 4.5 · high[10] | high | Mini-SWE-agent | 76.8% | $0.75 | — | $5/$25 |
| Sonnet 4.5 · high[10] | high | Mini-SWE-agent | 71.4% | $0.66 | — | $3/$15 |
| Haiku 4.5 · high[10] | high | Mini-SWE-agent | 66.6% | $0.33 | — | — |
| DeepSeek V4 Flash 0731[10] | max | Mini-SWE-agent | 88.8% | $0.01 | — | — |
According to Figure 4, the scatter plot compares the cost and capability of systems a team might use. Higher values on the vertical axis indicate stronger capability. Lower values on the horizontal axis indicate lower suite cost. On Verified Mini-SWE, Claude Opus 5 leads near 97.0% resolve at about $1.29/task, while GPT-5.6 Sol sits at 96.2% and $1.15/task. After OpenAI’s July 2026 list-price cut, GPT-5.6 Luna runs about $0.04/task at 93.0% resolve—roughly −97% less expensive than Sol on the same harness—with DeepSeek V4 Flash 0731 near 88.8% at about $0.01/task.[44][47][48] On DeepSWE, scores fall and costs spread further apart. More difficult work separates a strong headline score from a package inexpensive enough to run as default automation.
The spread of results makes cost central to model selection. Teams need to choose a model, reasoning effort, and harness that remain affordable at their daily workload. Teams that pursue headline resolve percentage alone overpay on high-volume loops and underfund the human review and design work those loops create. That reallocation of human hours is the irreversible move. The frontier advanced and laboratory volume expanded, yet model and harness selection still determine whether the leverage is economically sustainable.[10][43] Once that economics is established, organizational design must change with it. If the prior headcount model remains unchanged, the bottleneck simply relocates into review queues and coordination debt.
3 · Organizing for AI-native work
The economics in Figures 1–4 call for a different allocation of responsibility. AI-native teams make agents the default implementers and focus people on goals, verification, and harness design. Product ambitions can stay the same while the work needed to deliver them changes.
Coding is no longer the bottleneck.
According to Figure 5, the limited input flips from labor-as-typing to judgment-as-bottleneck. On the human-first side, headcount sets the calendar, projects wait on typing time, up-front process justifies itself when builds are costly, and review load rises only as people write more. On the agent-first side, agents draft and retry in parallel, verification and coherence set the calendar, evaluation suites and harnesses improve later runs, and review tracks agent throughput unless the organization redesigns paths and roles.
The operating model follows that economics. AI-native teams treat agents as default builders and keep humans on intent, verification, and harness design.[1] Working software takes precedence over ceremony, overnight agents arrive with harder review, and smaller cores own more surface area.[1][2] When overnight agents open a large volume of pull requests, the review queue sets the calendar rather than typing time. Teams that convert failures into harness rules retain the leverage;[4] teams that only add authors recreate the prior bottleneck in a new form.
The pattern is leaving the laboratories. Anthropic and NEC placed Claude Code in front of tens of thousands of employees.[23] Roles such as loop designers, evaluation owners, and fleet operators matter more than ceremonial planning.[22] Those roles justify their existence by designing how work is discovered, dispatched, and verified at scale. Loop engineering addresses that design problem by separating the work agents accelerate from the decisions people still own.
4 · Engineering the loop
Loop engineering organizes repeated work around clear goals and feedback. Product development has three connected rhythms, covering agentic coding, developer feedback, and external feedback. Each needs its own investment. Agents accelerate the inner coding loop; people still need systems for directing that work and deciding whether to accept the result.
Loop engineering is replacing yourself as the person who prompts the agent. You design the system that does it instead.
A loop is a recursive, goal-directed system that discovers work, dispatches agents, verifies results, records state, and repeats.[24] By mid-2026 that shift was concrete from three directions at once. Boris Cherny, creator of Claude Code, frames the work as writing loops that prompt Claude rather than prompting turn by turn.[25] Peter Steinberger, creator of OpenClaw, makes the same case for high-throughput open source by designing loops that prompt agents.[26] Andrew Ng, founder of DeepLearning.AI and co-founder of Coursera, elevates loop engineering to product cadence with three development loops at minutes, hours, and days.[27][28][29] The skill moves from prompt craft to system design. Ng’s three product-development loops place those cadences on a single map.
According to Figure 6, product work proceeds on three nested cadences. Agentic coding takes minutes, developer feedback takes hours, and external feedback takes days, with product specifications and developer vision as shared nodes. Agents accelerate only the innermost loop. Product decisions and market feedback still determine the pace of the outer loops.
Putting these rhythms into practice requires tools for scheduling, isolation, feedback, and persistent state. Across
Codex-style applications and Claude Code, the shared stack includes scheduled automations and goals, git worktrees for
parallel isolation, Skills (SKILL.md) as progressive project knowledge, MCP and connectors for real tools,
sub-agents so maker ≠ checker, and durable state on disk in the form of markdown, tickets, and AGENTS.md, because
models forget.[24] A typical daily loop is the short pipeline that turns those primitives into a closed
path from triage to ship.
According to Figure 7, the daily loop follows Triage → Isolate → Draft → Verify → Ship, with the Human responsible for exceptions and merging. Clear stop conditions keep unattended work within review capacity. Loops amplify the judgment built into them. Superficial review allows comprehension debt to accumulate.[30]
A person must still own the merge decision. Reliable execution also depends on the surrounding documentation, tools, tests, hooks, and memory. Together, these form the harness that makes a model useful as part of a trusted agent fleet.
5 · Engineering the harness
Agent = Model + harness. The harness supplies instructions, checks, tools, and persistent repository knowledge. It turns individual prompts into repeatable work and gives teams a way to carry lessons from one failure into later runs.
Agent = Model + Harness
A harness provides the model’s persistent operating environment through the following components.[4][22][31]
- System prompts,
CLAUDE.md,AGENTS.md, skill files, and subagent prompts - Tools, skills, MCP servers, and their descriptions
- Bundled infrastructure such as filesystem, sandbox, and browser
- Orchestration logic such as subagent spawning, handoffs, and model routing
- Hooks and middleware for deterministic execution, including compaction, continuation, and lint checks
- Observability such as logs, traces, and cost and latency metering
According to Figure 8, people direct the agent through guides and sensors. Guides set expectations before action; sensors provide feedback during execution. State stored outside the context window allows lessons from failures to become lasting harness rules. The harness turns a local correction into an improvement that later runs can reuse.
OpenAI’s experiment treats repository knowledge as the system of record. A short AGENTS.md is the table of contents,
approximately 100 lines, not the encyclopedia.[4] Structured docs/ holds design documents,
execution plans, product specifications, references, and quality grades. Progressive disclosure keeps the map in context
and the depth on disk. Per-worktree bootable applications and strict linters whose errors include remediation
instructions make the agent runnable without a human at every step.
Multi-hour agent runs require the same structure at longer horizons, with clear stop conditions, durable state, and feedback that keeps unattended work on track.[32] Conference practice adds hooks and a ratchet discipline.[2] Each observed failure becomes a permanent harness rule.[31] The map of what lives in context versus what lives on disk is then a cost and reliability problem as much as a documentation problem.
According to Figure 9, short AGENTS.md is the map injected into context, and structured docs/ is the encyclopedia on
disk. Progressive disclosure means the agent opens only the branch it needs.
That layout also determines what the model re-reads on every turn. System prompts, tool descriptions, skill files, and long agent maps enter the context window before the task does. If the full wrap is dumped at every step, token consumption rises rapidly, useful task context is displaced, and the prompt prefix changes enough that platforms cannot reuse work. A production harness designs that load deliberately.
- Short
AGENTS.mdmaps into progressivedocs/. The agent begins with a table of contents and opens depth only when the task requires it. - Keep system prompts and tool schemas byte-identical across turns. A stable prefix is what platforms can cache; churn reduces hit rates.
- Shrink live history with compaction and continuation hooks. The window holds the working set. Repository knowledge remains on disk as the durable system of record.
- Keep unused tool schemas and raw tool dumps out of the window. Tool definitions and bulky outputs displace the task long before they help.
- Inspect tool results programmatically and re-enter only required fields. One conference example reduced token use by about 66% by converting tool results from JSON to markdown and dropping unused fields.
- Treat prompts, tools, hooks, and metering as affordability levers. They keep multi-step agent loops viable when retries are the normal path.
Context efficiency directly affects the cost of the harness.
At Code with Claude 2026, Brad Abrams of the Claude Platform stated a production floor for that meter. Aim for at least an 80% prompt-cache hit rate before other agent optimizations. Mature product harnesses already clear a higher bar. Cursor, Replit, Perplexity, and Claude Code itself report cache rates in the 90s.[2] Cached tokens are less expensive and faster, and they do not count against rate limits. High hit rates indicate that the wrap is stable and lean enough for strong models and frequent agent retries to remain the default production path. The practical task is to allocate human time to the loops, harnesses, and verification that sustain these gains.
6 · Putting lower execution costs to work
The lasting advantage comes from investing in loops, harnesses, and verification capacity. The following practices apply that principle to daily work. Conference reports and laboratory adoption support the direction; each team still needs to assess what fits its codebase.
Workflow
- Route recurring work through agents before assigning it to humans. Triage, documentation, patches, and summaries are sound default tasks, while people retain responsibility for exceptions and the final merge decision.[1]
- Reserve real calendar time to understand code that no one on the team wrote by hand. A cursory review leaves the team without an adequate mental model of the code, allowing comprehension debt to compound.[30]
Harness and loops
- Keep agent maps short and use them as navigational indexes. Point from a concise table of contents to progressive detail in
docs/, keeping the active context inexpensive and current.[4] - Separate maker and checker roles, and define completion criteria before a loop begins. An unattended run needs an explicit condition that ends the work or returns it for human judgment.[24]
Organization and cost
- Resolve technical debates with working prototypes. When agents make experiments inexpensive, working software supplies decision evidence more quickly than additional process ceremony.[1]
Code is cheap, show me the Harness and Loop
Conclusion
Agents have reduced the cost of drafting, running, and repairing code. Human time increasingly goes to judgment, verification, and the harnesses that support reliable execution. Engineers need time to review changes, design systems, and coordinate agents. They also need time to understand the generated code they become responsible for maintaining.
The evidence offers a direction for investment. Output measures can overstate gains in judgment, and laboratory practice advances faster than adoption in regulated or legacy environments. Measure results in your own team, then invest in loops, harnesses, and verification.
Spend inexpensive execution on better decisions.
References
- Fiona Fung (). “Running an AI-native engineering org”. Code with Claude 2026 (YouTube). Argues that the bottleneck moved from coding to the work surrounding coding. Describes AI-native engineering practice, including prioritization of working software over process ceremony.
- Chris Ebert (). “Notes from Code with Claude 2026”. chrisebert.net. Conference notes from Code with Claude 2026 covering Fung’s bottleneck diagnosis, Abrams’s floor of at least 80% prompt-cache hit rate, mature product cache rates in the 90s, hooks as agent feedback, and overnight agent operation.
- Fiona Fung (). “Interview on Lenny’s Podcast”. Lenny’s Podcast / YouTube. Organizational diagnosis that coding is no longer the bottleneck for AI-native engineering teams.
- Ryan Lopopolo, OpenAI (). “Harness engineering: leveraging Codex in an agent-first world”. OpenAI. Reports an internal product of approximately one million lines of code with effectively no hand-written source. Describes AGENTS.md as an approximately 100-line table of contents, structured docs/ progressive disclosure, and a division of labor in which humans steer while agents execute.
- Addy Osmani, Director of Google Cloud AI (). “The 80% problem in agentic coding”. addyo.substack.com. Argues that review becomes the critical path as pull-request volume and size rise under AI adoption.
- Anthropic (). “2026 Agentic Coding Trends Report”. Anthropic resources. Industry report on the shift from writing code to reviewing and directing work under agent adoption.
- Sonar (). “How much time do developers spend actually writing code?”. sonarsource.com. Historical estimate of coding-time share (approximately 32%) prior to agent reweight of engineering attention.
- Software.com / Antenna (). “Global code-time sample”. antenna.dev. Independent sample of active coding time relative to the full developer work week.
- Leadership Garden (Faros / DORA-style synthesis) (). “AI: the work moved”. leadership.garden. Synthesis of industry telemetry: higher pull-request volume, longer review cycles, and larger diffs under high AI adoption.
- Vals AI (). “SWE-bench Verified leaderboard”. vals.ai. Shared Mini-SWE harness scores and cost per test for cross-model comparison on SWE-bench Verified.
- Anthropic Engineering (). “Raising the bar on SWE-bench Verified with Claude 3.5 Sonnet”. anthropic.com. Reports upgraded Claude 3.5 Sonnet at 49% on SWE-bench Verified (October 2024 model). Comparison table includes Claude 3 Opus at 22% and prior state of the art near 45%.
- OpenAI (). “OpenAI o1 System Card”. openai.com. Reports o1-preview at 41.3% on SWE-bench Verified under an Agentless scaffold.
- Anthropic (). “Claude 3.7 Sonnet and Claude Code”. anthropic.com. Reports Claude 3.7 Sonnet at 70.3% vendor SWE-bench Verified under a high-compute scaffold, and 63.7% with a simpler agent.
- Anthropic (). “Introducing Claude 4”. anthropic.com. Reports Claude Sonnet 4 at 72.7% on SWE-bench Verified at launch.
- OpenAI (). “Introducing GPT-5 for developers”. openai.com. Reports GPT-5 at 74.9% on SWE-bench Verified, with API list pricing of $1.25 and $10 per million tokens.
- Anthropic (). “Introducing Claude Opus 4.5”. anthropic.com. Reports vendor SWE-bench Verified of 80.9%, with list pricing of $5 and $25 per million input and output tokens.
- OpenAI (). “Why we no longer evaluate SWE-bench Verified”. openai.com. Explains why OpenAI no longer evaluates SWE-bench Verified and discusses limits of fair cross-vendor comparison on that yardstick.
- OpenAI (). “API pricing”. platform.openai.com. List prices for OpenAI API models used in suite-cost comparisons.
- Anthropic Institute (). “When AI builds itself”. Anthropic. Reports production metrics of more than 80% Claude-authored merges and approximately eightfold merge volume per engineer relative to 2024. Notes that line-of-code multipliers overstate true productivity gains.
- Open Source For You (recap of AI.cc / AICC report) (). “Enterprise AI costs crash 67% as open source models and multi-model routing go mainstream”. Open Source For You. Platform sample of blended enterprise token cost declining from $18.40 to $6.07 per million tokens between Q1 2025 and Q1 2026, a reduction of approximately 67%.
- METR (). “Time horizons”. metr.org. Measures of agent task time horizons roughly doubling on a multi-month cycle.
- Martin Fowler and Birgitta Böckeler (). “Harness engineering for coding agent users”. martinfowler.com. Distinguishes feedforward guides from feedback sensors and treats the harness as leverage when agents produce code.
- Anthropic (). “Anthropic and NEC”. anthropic.com. Announces Claude and Claude Code deployment for approximately 30,000 NEC Group employees as a large-scale AI-native engineering initiative.
- Addy Osmani, Director of Google Cloud AI (). “Loop engineering”. addyosmani.com. Defines loop engineering as the design of systems that prompt agents. Catalogues runtime primitives including worktrees, skills, and sub-agents with maker-and-checker separation.
- Boris Cherny (via circulated clip) (). “On writing loops that prompt Claude”. X / social synthesis. Creator of Claude Code: the work is to write loops that prompt Claude, rather than to prompt turn by turn.
- Peter Steinberger (). “Design loops that prompt your agents”. X. Creator of OpenClaw: for high-throughput development, design loops that prompt agents rather than prompting agents directly for each step.
- Andrew Ng (). “Loop engineering letter with three-loops diagram”. X. Public post of the three product-development loops diagram used as Figure 6 in this research topic.
- Andrew Ng (). “3 Key Loops for Building 0-to-1 Products with AI Agents”. LinkedIn / The Batch letter. Presents three product-development loops—agentic coding, developer feedback, and external feedback—at cadences of minutes, hours, and days.
- Andrew Ng (). “Three loops explainer”. YouTube. Video explanation of the three nested product-development loops and their respective cadences.
- Addy Osmani, Director of Google Cloud AI (). “Comprehension debt”. addyosmani.com. Risk of shipping code that no team member understands when agents author at scale.
- Addy Osmani, Director of Google Cloud AI (). “Agent harness engineering”. addyosmani.com. Advocates a ratchet discipline in which each observed failure becomes a permanent harness rule.
- Anthropic Engineering (). “Effective harnesses for long-running agents”. anthropic.com. Harness design for multi-hour agent runs, including structure, feedback, and stop conditions.
- Anthropic (). “Introducing Claude Sonnet 4.6”. anthropic.com. Coding upgrade for Sonnet 4.6 with standard list pricing of $3 and $15 per million tokens.
- Anthropic (). “Introducing Claude Sonnet 5”. anthropic.com. Agentic Sonnet tier with cost-performance curves across effort levels.
- Anthropic (). “Introducing Claude Opus 4.8”. anthropic.com. Frontier coding model; Vals Mini-SWE results in the 88.6% class.
- OpenAI (). “Introducing GPT-5.2”. openai.com. Reports GPT-5.2 Thinking at 80.0% vendor SWE-bench Verified.
- OpenAI (). “Introducing GPT-5.5”. openai.com. GPT-5.5 generation with scores in the Vals Mini-SWE period.
- OpenAI (). “GPT-5.6”. openai.com. GPT-5.6 Sol, Terra, and Luna family. July 9, 2026 launch list prices later updated July 30: Luna $0.20/$1.20, Terra $2/$12, Sol $5/$30 per million tokens.
- Moonshot AI (). “Kimi K2.7 Code”. kimi.com. Open coding-focused agentic model; Vals Mini-SWE Verified 78.2% at approximately $0.27 per test.
- Moonshot AI (). “Kimi K3: Open Frontier Intelligence”. kimi.com. Flagship open model; Vals Mini-SWE Verified 93.4%. Cost per test not yet published at the cited snapshot.
- Datacurve (). “DeepSWE leaderboard”. deepswe.datacurve.ai. Long-horizon software-engineering benchmark of 113 tasks under a mini-swe-agent harness, reporting pass@1 and average cost with contamination controls.
- OpenAI (). “Advancing the price-performance frontier with GPT-5.6”. openai.com. API list prices: GPT-5.6 Luna $0.20 input / $1.20 output per million tokens (80% reduction); Terra $2 / $12 (20% reduction); Sol unchanged at $5 / $30.
- OpenAI (). “GPT-5.6 Luna model card”. developers.openai.com. Documents Luna text-token list pricing of $0.20 per million input and $1.20 per million output after the July 30, 2026 price update.
- OpenAI (). “API pricing”. developers.openai.com. Flagship table lists gpt-5.6-luna at $0.20 input and $1.20 output per million tokens (short context).
- Anthropic (). “Introducing Claude Opus 5”. anthropic.com. Opus 5 launch at $5 / $25 per million tokens (same as Opus 4.8). Strong coding and agentic results on Frontier-Bench and related evals; Vals Mini-SWE reports 97.0% at about $1.29 per test.
- DeepSeek-AI (). “DeepSeek-V4-Flash-0731”. Hugging Face / deepseek-ai. Official Flash 0731 release superseding preview. Model card reports DeepSWE 54.4% among agentic evals; Vals Mini-SWE lists 88.8% resolve at about $0.01 per test.