Whether the enterprise agent-product surface — the layer where labs and adjacent vendors package existing model capability into something a company actually deploys (seats, IDE integrations, workflow bundles) — keeps shipping at the pace 2026-08-20 revealed, and whether this map can keep up now that it has a place to log it. Opened after the heaviest single-day coverage-critic miss on record: five confirmed misses in one day (08-20), all on this exact axis, zero overlap with anything this map tracked, and only one (Mistral's Agentic Search) had an existing thread to land on. The candidate this produced was offered twice (08-20, 08-21) and dropped without a decision either way — a real, evidence-backed gap that sat unresolved for five days before being promoted. Track: whether the pace holds or 08-20 was a one-off pileup; whether any lab treats this as a distinct product line with its own roadmap rather than a bundling exercise; and whether enterprise adoption numbers (seats, IDE install counts) ever get disclosed to test whether the packaging actually converts to usage.
Sam Altman publicly apologized for the Astra launch’s staged access, which prioritized OpenAI’s own Daybreak cybersecurity-tester cohort and left paying ChatGPT subscribers — including Pro subscribers, who normally get first access to new releases — watching enterprise customers go first instead. In his own words on X: “first, sorry for the messy rollout. second, when we screw up, we try to make it right. third, we should be able to begin broad rollout to API customers and chatgpt subscribers in the near future. as usual we will start with pro subscribers.” OpenAI is compensating with “one banked reset for every day you don’t have access to Astra on your paid ChatGPT plan,” and Altman said he was hopeful (not promising) Pro subscribers could use it over the 09-05/09-06 weekend. This directly complicates this thread’s own 09-03 entry, which read Astra reaching “Business and Enterprise tiers alongside Plus and Pro within days” — the “within days” promise ran into the same Daybreak-first gating that Frontier Gatekeeping already tracks as this thread’s lab-run access-control pattern. (Sam Altman on X, The Verge, Unite.AI)
OpenAI silently changed several of GPT-6 Astra’s headline benchmark figures in the hours and days after its own launch post went up, Fortune reported, rather than the independent-evaluator gap this thread already has on record. Astra’s stated hallucination rate went from 4.2% (in the version live at 2:23pm ET on launch day) to 2% roughly three hours later, then back to 4.2%. On ExploitBench, the score OpenAI published for predecessor GPT-5.6 Sol changed from 5.5% to 11.5%, which OpenAI told Fortune reflects “a reasoning level that is not commercially available” for Sol — i.e., the comparison model was tested with a configuration customers cannot actually buy. On ARC-AGI-3, an embargoed draft carried 98.6%, the live blog post read 99.99%, and the Arc Prize Foundation’s own “powerful harness” test scored it 99.9% against 63% on the standard harness — explaining a discrepancy the 09-04 digest’s coverage-critic flagged and left unresolved (two AI newsletters had quoted a 99.9%/99.99% figure matching neither the 98.6% headline nor the then-known 62.7% standard-harness number). OpenAI’s on-record response: “we care deeply about getting evaluations right. Most evaluations have noise within a few percentage points based on the exact checkpoint, scaffold, and eval run used in reporting,” calling draft-to-final adjustment standard practice. Some outside researchers call it “benchmaxxing.” This is a new axis on top of the Artificial Analysis/ARC Prize dispute already logged 09-04 — not that outsiders measured differently, but that OpenAI’s own published numbers moved after the fact, and in one case a score for its own predecessor model rose on a configuration customers cannot buy. (Fortune, Startup Fortune)
Artificial Analysis published version 4.2 of its Intelligence Index on 09-04, an interim update it said it had “deliberately held back” through the week’s launches but released because “the frontier moving so quickly in the past weeks” made it necessary: AA-Briefcase (a private agentic knowledge-work set) and Surge AI’s GDP.pdf are added, the saturated GPQA Diamond is dropped, and 40% of the weighting now sits on private held-out sets, double v4.1’s. On the re-weighted index Claude Fable 5.1 leads, GPT-6 Astra is second with a four-point gain over GPT-5.6 Sol (it had scored level with Sol on v4.1), and Meta is the third-ranked lab ahead of SpaceXAI, Moonshot, Z.AI and Google. AA’s own note does not mention Fortune’s report that OpenAI changed Astra’s published figures after launch; The Decoder’s reading that the overhaul was “likely in response to criticism” of the Astra scoring is The Decoder’s. For the benchmark dispute this thread carries, the third-party index whose flat Astra score fed the skepticism now shows a gain — on a different test set, which is the point. (Artificial Analysis, The Decoder)
OpenAI’s launch materials call Astra “the world’s best computer use model,” citing an OSWorld 2.0 offline-subset score of 72.6% completed in roughly 40 minutes per task, against predecessor GPT-5.6 Sol’s 65.7% at roughly 75 minutes per task — a claimed 47% completion-time improvement — for a model built to navigate browsers, spreadsheets, websites and desktop applications and carry out multistep agentic workflows. It reaches ChatGPT’s Business and Enterprise tiers alongside Plus and Pro within days of the Thursday cybersecurity-tester-only start, plus the OpenAI API, AWS Bedrock and Microsoft Azure — priced at $10/million input and $50/million output tokens standard ($20/$100 in a 2.5x-faster “Fast” mode). (TechCrunch, VentureBeat)
OpenAI published a 39-page paper, “Improved Short Gaps Between Primes,” proving that infinitely many consecutive primes differ by at most 186 — the paper’s own introduction cites Polymath 8b’s 246 from 2014 as the bound it builds past — and stating in its abstract that “the proof is due to GPT 6 Astra,” with a machine-checked Lean formalization, as part of the 09-03 Astra launch materials. The method combines Polymath 8a and Stadlmann’s equidistribution estimates with new factorization conditions that enlarge the support of the multidimensional Selberg sieve, establishing DHL[40,2]; mathematicians Weijie Su and Perry Metzger noted the result the same day. The third lab pure-math claim in one week alongside Anthropic’s Fermat formalization, and the second launch-day Astra claim this thread carries whose checking is left to outsiders. (OpenAI — paper PDF, Weijie Su on X)
Salesforce closed up 22.58% at $252.05, its strongest single day since 2020, on a Q2 FY2027 beat announced with “Claudeforce” — Claude as the reasoning model behind the Atlas Reasoning Engine and Agentforce, the default model across Slack, and a “Salesforce in Claude” plugin carrying 37 prebuilt sales skills, in open beta expected September. Revenue was $11.345bn, up ~10.8% YoY, adjusted EPS $5.90 against ~$3.27 expected, FY2027 guidance raised $300M to $46.1-46.4bn. The land-grab reading is distribution: a frontier lab reaching enterprise seats through an incumbent’s install base rather than its own product surface. ⚠️ No dollar value is attached to the partnership in either company’s materials; “Claudeforce” is Salesforce’s own branding and does not appear on Anthropic’s newsroom, though Amodei is quoted in Salesforce’s release; and $200M of the $300M guidance raise is attributed to the pending Contentful and Fin acquisitions, not organic growth. (Salesforce investor relations)
Anthropic opened a research preview of the Model Hardware Standard (MHS) on 08-27 — a model-agnostic specification for AI agents to operate physical devices with a programmable interface (microscopes, liquid handlers, robotic arms), developed with HHMI Janelia Research Campus and shared first with scientific labs and advanced manufacturers “ahead of making the standard open source.” Anthropic’s own claim is that MHS cuts hardware integration from weeks or months to hours or minutes and lets agents run round-the-clock experiments, updating parameters in real time and in some cases recovering from hardware errors unaided; the MCP playbook — a shared protocol that becomes the default layer — extended from software tools to physical equipment. Both The Rundown AI and TLDR AI led their 08-28 editions with it, and this map had nothing — the clearest double-benchmark miss of the week, on the thread whose question is exactly where the lab-owned product surface goes next. ⚠️ Research preview with named partners, not a shipped product; no adoption number. (Anthropic, The Rundown AI)







