AI Research Story 1 of 12
Anthropic Publishes the First Hard Numbers on How Much of Its Own Research AI Now Leads
Anthropic has done something the frontier labs have so far only gestured at. It put a number on how much of its own artificial intelligence research and development is now led by its own model, and then published the measurement methodology alongside it.
The headline figure is 26 percent. As of August 2026, Claude leads 26 percent of Anthropic's AI R&D work, meaning the model completes most of a task end to end from a high level prompt with a human supervising rather than directing. Anthropic's report puts the same measure below 1 percent in February. That is a move from rounding error to a quarter of the work in roughly half a year, inside the company with the clearest view of what its own model can actually do.
The framework behind the number matters as much as the number. Anthropic scores work on an autonomy ladder running from no AI involvement at the bottom to full autonomy at the top. The rung below "leads" is "collaborates," where the model handles large chunks of work under close human direction. Measured that way, the share of Anthropic's AI R&D work sitting at or above "collaborates" is above 90 percent. The interesting boundary is no longer whether the model participates. It is how far up the ladder each task has climbed.
Scale gives the figure its texture. Anthropic reported roughly 30,000 agents running concurrently on its primary internal platform. This is not a handful of researchers experimenting with a chat window. It is a standing population of software processes doing research work, monitored, logged and measured, with a governance layer wrapped around it.
The company also disclosed how it splits its research compute. About 6 percent of the compute that went to AI R&D was allocated toward safety. Executives reading this should notice that the disclosure exists at all. A safety compute ratio is the kind of metric that becomes a benchmark the moment one company publishes it and a competitor does not.
For anyone running a technology organization outside a frontier lab, the useful read is not the 26 percent itself. It is the shape of the curve and the fact that someone is now measuring it publicly. Most enterprises have no equivalent instrument. They know how many seats of a coding assistant they bought and roughly what their engineers say about it. They do not know what share of their engineering work the model is leading rather than assisting, because nobody defined the ladder or scored the tasks against it.
Anthropic has now supplied both the ladder and a worked example. The obvious next question for a board is what the same measurement would return inside its own company, and whether anyone can currently answer.
AnthropicAutomationAI ResearchGovernance
Enterprise AI Story 2 of 12
OpenAI Ships Its First Vertical Edition of GPT-6 Astra, and It Goes to the Legal Market
OpenAI launched Astra for Law, a configuration of its flagship GPT-6 Astra model built specifically for legal research and drafting. It is the first time the company has taken its newest frontier model and shipped a named vertical edition of it, which makes the choice of vertical worth reading closely.
The product is not a prompt template. OpenAI paired the model with a dedicated legal search index covering more than 230 million URLs of United States case law, statutes, regulations, court rules and administrative decisions, plus instructions tuned for legal analysis and writing and a set of governance controls aimed at firm requirements.
The accuracy claim is where the commercial argument lives. On a set of 200 United States legal research questions, OpenAI reported that Astra for Law passed correctness checks on 54.0 percent, against 38.7 percent for GPT-6 Astra using ordinary web search. That is a substantial relative gain, and it is also an admission worth sitting with. On the company's own evaluation, its best general model with web search got fewer than two in five legal research questions right. Retrieval quality, not raw model capability, is doing most of the work here.
Distribution reflects a two sided strategy. OpenAI named Harvey and Legora among the API customers building on Astra for Law, which means the company is simultaneously selling to law firms and supplying the legal technology vendors who sell to the same firms. That is a deliberately uncomfortable position to occupy, and it is the position cloud providers have occupied for years.
For general counsel and chief information officers, three questions follow. The first is whether a 54 percent pass rate on research questions is good enough to change a workflow, and the honest answer depends entirely on what sits downstream. A researcher who checks every citation gains speed. A workflow that trusts the output gains liability.
The second is procurement. If the legal technology vendor you already pay is building on the same model you could license directly, you now have leverage in a renewal conversation you did not have last week, and your vendor now has to explain what it adds on top.
The third is the pattern. A frontier lab that ships one vertical edition rarely stops at one. Legal is a market with high billing rates, heavy document volume and clear evaluation criteria, which makes it a natural first target. Financial services, healthcare and accounting share those traits. Any company whose software vendors sit between it and a general model should assume the gap is narrowing and plan accordingly.
OpenAILegal TechGPT-6 AstraEnterprise AI
Policy & Regulation Story 3 of 12
The House Votes 417 to 3 to Push Data Center Grid Costs Off Household Bills
The United States House of Representatives passed H.R. 9340, the Ratepayer Protection Act, by 417 votes to 3 on September 16. In a Congress that agrees on almost nothing, a margin like that is the story. Data center electricity costs have become one of the few issues where the political incentive runs in a single direction.
The mechanism is more careful than the headline suggests, and executives should understand the difference. The bill does not impose a federal rule on how utilities bill large customers. It requires state regulatory authorities and nonregulated utilities to consider adopting standards under which large load customers of 100 megawatts or more cover the full incremental cost of the generation, transmission and distribution upgrades needed to serve them. States get one year to begin considering the standard and two years to reach a decision.
That is a consideration requirement rather than a mandate, a structure federal energy law has used before. It does not guarantee any particular outcome in any particular state. What it does guarantee is a proceeding, in every state, on a politically charged question, with a clock attached. For a company planning a large load interconnection, that is a scheduling problem as much as a cost problem.
The substantive detail that deserves attention is the inclusion of generation. Interconnection agreements have traditionally covered transmission and distribution, the wires and substations needed to deliver power. Extending the cost responsibility to the power supply itself is a meaningfully broader ask, and it lands on exactly the category of customer whose demand is growing fastest.
The economics of a large campus change if the developer carries the full incremental cost of new generation rather than sharing it across a rate base. So does the site selection calculus. States that move early and clearly may become more attractive than states that leave the question open, because certainty has value even when the answer is expensive. States that adopt aggressive standards may find projects routed elsewhere, which is the tension every commission will now have to resolve in public.
The bill still needs the Senate, and a 417 to 3 House vote is not a prediction of what the Senate does. But the political signal is already delivered. Residential electricity prices are now a live national issue, data centers are the named cause, and a near unanimous chamber has voted to make the industry pay its own way.
Any capital plan that assumes today's cost allocation rules survive the next two years is carrying a risk that just became explicit. The prudent move is to model the full incremental cost case now, before a state commission models it for you.
Energy PolicyData CentersRegulationUtilities
Funding & Investment Story 4 of 12
Crusoe Raises $3.9 Billion at a $30.9 Billion Valuation and Bets on Trucking Data Centers to the Power
Crusoe announced a $3.9 billion Series F at a $30.9 billion post money valuation, jointly led by Atreides Management, Mubadala Capital and Valor Equity Partners. The round is roughly three times the size of the company's previous raise and it repriced the business dramatically.
The comparison is the point. In October 2025 Crusoe announced the initial closing of an anticipated $1.375 billion Series E that brought its expected valuation to over $10 billion. Less than a year later the valuation has roughly tripled. Few categories in the current market have repriced that fast, and the ones that have are almost all sitting on the same bottleneck.
The bottleneck is power and time. The scarce resource in artificial intelligence infrastructure is no longer chips alone. It is an energized site with sufficient interconnection capacity, and the multiyear queue between deciding to build and being able to switch something on. Crusoe's answer is a modular product line called Crusoe Spark, which the company says shortens data center construction in the field from years to weeks by manufacturing the units rather than building them on site.
That framing inverts the traditional model. A conventional hyperscale campus takes years because the building is constructed where the power is, at whatever pace permitting, supply chains and labor allow. A factory built module can be produced continuously and delivered to wherever capacity happens to exist, including smaller pockets of stranded or underused generation that could never justify a full campus. It turns a construction problem into a manufacturing and logistics problem, which is a problem the industrial economy already knows how to scale.
Whether the economics hold is the open question. Modular construction trades site specific optimization for speed and repeatability, and it introduces transport, siting and servicing constraints that a single large campus does not have. Investors clearly believe the trade is worth making at this valuation.
For enterprise buyers the practical implication is about where compute becomes available and how quickly. If modular deployment works at scale, capacity stops being concentrated only in the handful of regions with enormous interconnection capacity and starts appearing closer to where power is available and where data needs to stay. That has consequences for latency planning, for data residency, and for anyone whose capacity roadmap assumes the current geography of availability is fixed.
The investor list is also a signal in itself. Sovereign wealth capital, infrastructure specialists and a chip maker all took positions in the same round. That is the profile of a company being funded as infrastructure rather than as software, with the duration expectations that implies.
CrusoeData CentersVenture CapitalAI Infrastructure
AI Infrastructure Story 5 of 12
Amazon Takes a Warrant in Generac to Lock Up Backup Power for Its Data Centers
Generac entered a long term supply agreement with Amazon under which a warrant vests against potential aggregate gross payments of up to $8 billion for backup generators. The structure of the deal is more interesting than the number attached to it.
Generac issued a warrant for 1,693,745 shares at an exercise price of $200.9266 per share to an Amazon subsidiary. Generac expects about $2.4 billion of initial deliveries in 2027 and 2028. Generac shares rose as much as 40 percent after the agreement was disclosed, according to market reporting.
An equity linked supply agreement is not how backup generators have historically been purchased. The customer is taking a financial position in the supplier, with the warrant vesting as actual purchase volume accumulates. Both sides are signing up to a relationship rather than a purchase order. Amazon gets priority claim on constrained manufacturing capacity and a stake in the value its own demand creates. Generac gets committed volume it can plan factories around, and a marquee customer whose name repriced its stock in a session.
The structure exists because of scarcity. Backup generation for data centers has become a genuine supply constraint, with lead times measured in years rather than months. When capacity is that tight, a hyperscaler cannot simply buy its way to the front of the line with a higher price, because the constraint is physical rather than financial. It has to underwrite the expansion of the line itself. The warrant is the underwriting instrument.
Executives should read this as a template rather than as a one off. Any input where demand growth has outrun manufacturing capacity is a candidate for the same treatment, and the artificial intelligence buildout has produced a long list of them. Transformers, switchgear, turbines, cooling equipment and high voltage components all share the profile.
There is a competitive consequence for everyone who is not Amazon. When the largest buyers lock in capacity through equity linked multiyear commitments, the remaining supply available to ordinary buyers shrinks and gets more expensive. A mid sized enterprise planning its own facility, or a colocation provider serving one, is now negotiating for what is left after the hyperscalers have taken their positions.
The narrower lesson is about corporate development. A procurement problem became a capital markets transaction, structured by people who understood both. Companies whose procurement and treasury functions do not talk to each other will keep treating supply constraints as sourcing problems and will keep losing to buyers who treat them as financing problems.
AmazonGeneracSupply ChainData Centers
Policy & Regulation Story 6 of 12
Scotland Stops Deciding on Large Data Centers Until It Writes the Rules
On September 16 the Scottish Parliament passed amendments stating that no planning or consenting decisions on large data centre applications should be made until new national planning guidance is completed. The wording matters, because what Scotland did is being described in ways the vote does not support.
The Scottish Greens brought a motion calling for an outright moratorium. The governing parties did not want the word. Scottish Government minister Hannah Mary Goodlad said of data centres, "We do not agree that there should be a moratorium." What passed instead were amendments that suspend decisions rather than applications, which produces a pause in practice while avoiding a formal ban in law.
The Labour amendment that carried requires the Scottish Government to report on national planning guidance for data centres by the end of the calendar year and to publish that guidance in full within 12 months. Pending decisions include projects in Falkirk, North Lanarkshire, Inverclyde and Fife.
The distinction between a moratorium and a decision pause is not rhetorical. A moratorium tells developers to stop applying. A decision pause tells them to keep applying and wait, which preserves the pipeline and preserves the government's ability to say it remains open for investment. For a developer with capital committed and a construction schedule, the practical effect is similar in the near term and quite different in the medium term. The application does not die. It queues.
Scotland is an instructive case because its selling point and its problem are the same thing. Substantial renewable generation makes it attractive to operators who need clean power at scale. That same generation sits on a grid with constraints, and the local political question of who benefits from the electricity and who absorbs the consequences of large new loads has now been asked out loud in a parliament.
This is the second jurisdiction level intervention on data center siting within a week of the United States House voting on ratepayer cost allocation. The pattern across both is not opposition to artificial intelligence. It is a demand that the terms be written down before the capacity is approved. Regulators are discovering that they approved a generation of facilities under rules written for a different scale of electricity demand, and they are pausing to rewrite them.
For anyone with a European capacity roadmap, the planning assumption to revisit is timeline certainty. A jurisdiction that pauses decisions for up to 12 months while it writes guidance is not necessarily a hostile jurisdiction, but it is an unpredictable one during the drafting window. The operators who fare best will be those who engage with the guidance while it is being written rather than after.
ScotlandData CentersPlanningEnergy
AI Infrastructure Story 7 of 12
India Opens the Second Phase of Its Semiconductor Mission as SEMICON Comes to Delhi
Prime Minister Narendra Modi launched the second phase of the India Semiconductor Mission at SEMICON India 2026 in New Delhi. The event ran September 17 to 19 at Yashobhoomi in Dwarka, and the timing of the launch against the conference was not accidental. India wanted the announcement made in front of the industry it is trying to recruit.
The first phase provides the benchmark. The Semicon India Programme, also called Semicon 1.0, carried a total outlay of Rs 76,000 crore. As of December 2025 it had approved 10 projects with total investment of Rs 1.60 lakh crore, which means the incentive framework mobilized roughly twice its own value in committed private capital. Whatever else is arguable about industrial policy, that multiplier is the number the second phase is being sold on.
India's semiconductor strategy has been deliberately unglamorous. Rather than chasing leading edge logic fabrication, where capital intensity and technical risk are extreme and the incumbents are decades ahead, the program has concentrated on assembly, testing, marking and packaging, on design, and on the equipment and materials that feed the industry. That is a strategy of entering the value chain where entry is actually possible and building upward from there.
Advanced packaging in particular is a more interesting position than it sounds. As transistor scaling delivers diminishing returns, packaging has become one of the main sources of performance improvement in artificial intelligence accelerators. The techniques that stack and connect chiplets are now genuinely differentiating rather than commodity back end work. A country that builds real capability there is not on the periphery of the industry.
The strategic context is the obvious one. Semiconductor supply chains concentrated in a small number of geographies have become a standing concern for every government and every large buyer, and the answer everyone has settled on is geographic diversification even at a cost premium. India is competing for that diversification against the United States, Japan, Europe and Southeast Asia, all of which are offering incentives of their own.
For technology executives, the planning horizon is what matters. Fabrication and packaging capacity announced today produces output years from now, and the programs launched in this cycle will shape where components are made a decade from now. Procurement organizations building supplier diversification strategies should be tracking which of these programs actually converts commitments into operating plants, because announcements and capacity are not the same thing and the gap between them is where most industrial policy disappoints.
The measure to watch in India is not the size of the next incentive package. It is how many of the 10 approved projects reach volume production on schedule.
IndiaSemiconductorsIndustrial PolicySupply Chain
Enterprise AI Story 8 of 12
Novo Nordisk Puts Claude Science Inside Its Drug Discovery Pipeline
Novo Nordisk and Anthropic announced a collaboration to advance drug discovery with Claude on September 16. Novo will use Claude Science for biological reasoning in research and development workflows, alongside applying the models to strengthen its software engineering capability.
The names attached tell you the level at which this was decided. Mike Doustdar, President and Chief Executive Officer of Novo Nordisk, framed it in productivity terms, saying that "AI can help us increase productivity in R&D and compress the path from research to marketed product." Dario Amodei, co founder and Chief Executive Officer of Anthropic, made the broader claim that artificial intelligence's increasing capability "brings with it the potential to compress a century's worth of biological and medical breakthroughs into a decade."
Strip away the ambition and the specific mechanism is worth understanding. Pharmaceutical research and development is not primarily bottlenecked on laboratory throughput. It is bottlenecked on reading, synthesis and judgment. A research program lives inside an enormous and constantly growing literature, a deep internal archive of prior experiments including the failures, and a regulatory record that constrains what is worth attempting. The scarce resource is the ability to hold all of that in view while deciding what to try next.
That is a reasoning and retrieval problem before it is a chemistry problem, which is why a general frontier model configured for scientific work is a plausible tool rather than a marketing exercise. The second half of the collaboration, applying the models to software engineering, is the less discussed and possibly more immediately measurable half. Large pharmaceutical companies run substantial internal software estates for trial management, data pipelines and regulatory submission, and that work is closer to the current proven capability of these models than novel biology is.
The enterprise lesson generalizes past healthcare. The deployments that produce durable value are the ones aimed at a specific bottleneck that the organization can already name and measure, rather than broad enablement programs that distribute licenses and hope. Novo named two bottlenecks, put a product against each, and kept human oversight and data governance explicit in the announcement.
For boards evaluating their own programs, the diagnostic question is simple and uncomfortable. Can you state, in one sentence, which constrained step in your value chain a given artificial intelligence deployment is meant to relieve, and what you will measure to know whether it did? Most programs cannot answer. The ones that can tend to be the ones that survive the second budget cycle.
A pharmaceutical company that shortens its research timeline by even a modest margin changes the economics of its entire portfolio, which is why this category of deployment attracts chief executive attention rather than delegated sponsorship.
Novo NordiskAnthropicDrug DiscoveryLife Sciences
Industry Dynamics Story 9 of 12
Pew Finds the Politics of AI Anxiety Have Flipped
The Pew Research Center surveyed 3,488 United States adults between June 22 and June 28, 2026, and found something that upends a common assumption about where resistance to artificial intelligence comes from. Among Democrats, 56 percent said they are more concerned than excited about artificial intelligence in daily life, against 49 percent of Republicans.
The employment finding is sharper. Among Democrats, 75 percent expect artificial intelligence to lead to fewer jobs over the next 20 years, up 17 percentage points from 2024, against 68 percent of Republicans. A 17 point move in two years is not drift. It is a realignment, and it happened while the technology was being deployed into exactly the kinds of work that concentrate in the Democratic coalition.
That is the mechanism worth understanding. Earlier waves of automation fell hardest on manufacturing and logistics, which is why automation anxiety tracked one way politically for decades. The current wave lands first on writing, analysis, design, coding, customer support, marketing and administration. Those are urban, credentialed, salaried occupations, and they are distributed very differently across the electorate than assembly line work was.
For executives the immediate consequence is internal rather than electoral. Three quarters of one large segment of the workforce now expects fewer jobs as a consequence of this technology. That expectation shapes how any deployment is received inside a company, regardless of what the deployment is actually designed to do. An employee who believes the tool is a prelude to a reduction in force will use it differently, report on it differently, and talk about it differently than an employee who believes it removes work they dislike.
Companies that have not addressed this directly are running an adoption program on top of an unstated assumption that the program is a headcount exercise. Silence does not read as neutrality in that environment. It reads as confirmation.
The policy consequence follows the political one. When concern becomes bipartisan, with both major parties above two thirds on the job loss question, regulation stops being a partisan fight and starts being a question of which approach wins rather than whether anything passes. The near unanimous House vote on data center electricity costs this same week is what that looks like in practice.
The strategic read is that the window for shaping how artificial intelligence gets regulated in the United States is narrowing, and it is narrowing because public opinion consolidated faster than most corporate government affairs functions planned for. Companies that assumed they had years to influence the framework may find the framework arrives first, written by people responding to a constituency that has already made up its mind.
Public OpinionWorkforcePolicyPew Research
AI Business Models Story 10 of 12
Temporal Raises $550 Million on the Unglamorous Business of Making Agents Finish What They Start
Temporal raised a $550 million Series E at a $12.55 billion valuation, announced September 14. The company sells durable execution, which is infrastructure that guarantees a long running workflow survives the failure of any individual component and resumes where it left off rather than starting over or silently stopping.
The growth figures explain the valuation. Temporal's annualized revenue run rate is up more than 200 percent year over year. Temporal Cloud processed 1.9 trillion billable actions in August 2026, up more than 350 percent year over year. The company has more than 4,300 paying customers, up 139 percent year over year, and says OpenAI's use of Temporal has grown 60 fold in under a year.
That last figure is the one to sit with, because it describes what actually happens when agent systems move from demonstration to production. An agent that performs a single task in a single session needs almost none of this. An agent that runs for hours across a dozen external systems, waits on human approval, retries failed calls and has to be auditable afterward needs all of it. The hard part of that system is not the reasoning. It is the state management, and state management is a solved problem in a way that reasoning is not.
There is a useful pattern here for anyone evaluating where value accrues in this market. The models attract the attention and the capital, and they are also the layer most subject to substitution, price compression and rapid obsolescence. The layer underneath, which tracks what happened, guarantees it completes, and makes the whole thing debuggable, is stickier by construction. Once workflow state lives in a system, moving it is a migration project rather than a configuration change.
The 350 percent growth in billable actions against 139 percent growth in paying customers is the detail most worth reading carefully. Customers are growing, but usage per customer is growing considerably faster. That is the signature of deployments expanding from pilot to production rather than of a customer count inflated by trials. It is also the pattern that makes revenue durable, because usage that grows inside an existing account does not require winning a new sales cycle.
For enterprise architects the practical question is whether your own agent deployments currently have an answer to what happens when a long running process fails halfway through. Many do not, because the pilot never ran long enough or at enough volume for the question to become urgent. It becomes urgent precisely at the moment the deployment starts to matter, which is the worst possible time to discover the gap.
TemporalAgentsInfrastructureVenture Capital
Generative AI Story 11 of 12
Profound Raises $180 Million Betting That Brands Now Have to Market to the Model
Profound raised a $180 million Series D at a $1.8 billion valuation led by Sequoia Capital and Kleiner Perkins. The company sells visibility inside artificial intelligence generated answers, a category that did not meaningfully exist two years ago. Profound says about a third of the Fortune 100 use its platform.
The pace of repricing is the first thing to notice. Profound raised $96 million in a Series C at a $1 billion valuation before this round. The valuation nearly doubled in the interval, which tells you less about this particular company than about how quickly the market decided the underlying problem is real.
The problem is straightforward once stated. For roughly two decades, the discipline of getting found online meant optimizing for a ranked list of links that a person then chose from. Increasingly, the user never sees the list. They ask a question and receive a synthesized answer that names some products, companies and sources and omits all the others. Whatever governs that selection is now a commercially decisive variable, and almost no marketing organization has instrumentation pointed at it.
That is the gap Profound is selling into. The classic search optimization toolset measures rank, traffic and clicks. None of those describe whether a model recommends you when someone asks it to compare options in your category, and a brand can be entirely absent from that surface while its traditional search metrics look healthy. The first thing most companies buy here is simply the ability to see the problem.
There are real reasons for executives to be careful. The mechanisms by which models select what to mention are opaque, unstable across versions, and not under anyone's direct control. A vendor can measure how often a brand appears in generated answers with reasonable rigor. Claims about reliably changing that frequency deserve harder scrutiny, because the underlying system can change without notice and take any inferred technique with it.
The strategic point survives that caution. A growing share of the moment where a customer forms an impression of a category now happens inside a model's answer rather than on a page any company controls. Marketing organizations that do not measure their presence in that surface are flying on instruments calibrated for a channel that is losing share.
The prudent posture is to treat measurement as the near term purchase and influence as an experiment. Know where you stand in generated answers, track it over model versions, and be skeptical of anyone who promises to move the number the way an agency once promised to move a search rank.
ProfoundMarketingAI SearchVenture Capital
AI Models Story 12 of 12
A New Benchmark Rebuilds Whole Applications, and the Best Model Manages Half
A new benchmark called ProgramDistill tests something most coding evaluations avoid. Rather than asking a model to fix an isolated bug or implement a single function, it asks the model to reconstruct working applications, and it verifies the result against recorded behavior rather than against a description.
The construction is substantial. ProgramDistill comprises 4,063 tasks built from 26 applications with 1,975 replay verified behaviors. The verification approach is what makes it interesting. Rather than checking whether code passes tests a human wrote, it checks whether the rebuilt application actually behaves as the reference application behaved, which is a considerably harder and considerably more honest standard.
The results are the part executives should read. On cumulative workflows in full application reconstruction, GPT-6 Astra reached 49.2 percent and Claude Opus 5 reached 28.8 percent. The leading model manages roughly half. The gap between the two leading systems is also wide, which is itself informative in a market where frontier models are often discussed as interchangeable.
Both numbers deserve interpretation rather than celebration or dismissal. A model that can reconstruct half of a working application from reference behavior is doing something that would have been science fiction three years ago. A model that fails the other half is not a system you hand an application rewrite to and walk away from. Both statements are true simultaneously, and most organizational disappointment with these tools comes from acting on only one of them.
The word doing the work in the results is cumulative. Reconstructing an application is not one task. It is a long sequence of dependent tasks where an error early on propagates through everything after it. Single step benchmarks flatter models because each attempt starts from a clean state. Cumulative benchmarks measure what actually matters in production work, which is whether the system can stay coherent across a long chain of decisions without drifting.
That distinction maps directly onto enterprise experience. Teams consistently report that these models are excellent at bounded, well specified work and degrade as scope and duration grow. ProgramDistill puts a number on the degradation rather than leaving it as anecdote.
The practical guidance follows. Scope engineering work to these models in units where the success rate is high and verification is cheap, and maintain human ownership of the architecture and the sequencing. The organizations getting real leverage are not the ones handing over whole projects. They are the ones that have gotten good at decomposition, which is a management skill rather than a technical one, and which is the actual bottleneck in most enterprise deployments.
BenchmarksCodingGPT-6 AstraClaude Opus 5