AI Daily Digest

Thursday, June 18, 2026

5,431 words · All issues

Top items

  • SpaceX acquires Cursor (Anysphere) for $60B all-stock, days after its record IPO, with a new “generally intelligent” from-scratch model (Composer 3) in the works.
  • Z.ai releases GLM-5.2, an MIT-licensed open-weights model with a 1M-token context window that nears GPT-5.5 and Claude Opus 4.8 on coding benchmarks.
  • Leaked audited financials show OpenAI lost $38.5B in 2025 (revenue $13.07B), as ChatGPT’s market share slips below 50% for the first time — right as it files confidentially to IPO.
  • The Anthropic Fable/Mythos export-control saga deepens: the supposed “jailbreak” was literally the prompt “fix this code,” the US denied even a UK carveout, and rivals (DeepSeek’s $7.4B raise, Z.ai, Cohere, Mistral) are reaping the fallout.
  • NVIDIA Blackwell (GB300 NVL72) sweeps all seven MLPerf Training 6.0 benchmarks at 8,192-GPU scale; CoreWeave trained DeepSeek-V3 671B in ~2 minutes.
  • Microsoft’s Copilot Cowork and Databricks’ Genie One both go GA, intensifying the enterprise “agentic coworker” race.

Company & product developments

SpaceX acquires Cursor (Anysphere) for $60B in all-stock deal. SpaceX formally exercised its option — first announced in April — to buy the AI coding startup Cursor, parent company Anysphere, in a deal paid entirely in SpaceX stock and expected to close in Q3 subject to regulatory approval. The timing is central: SpaceX listed publicly on Friday at $135/share, rocketed past $200 (some reports say it nearly doubled in under a week), piling roughly $1T onto its value and pushing Elon Musk’s personal wealth above $1T — making a stock-only acquisition far easier. The April agreement had given SpaceX the choice of acquiring Cursor for $60B or paying $10B for a partnership alone. Cursor, founded in 2022 as Anysphere, went through OpenAI’s startup accelerator in 2024, raised about $3.38B from Thrive, a16z, the OpenAI Startup Fund and Nvidia, and crossed $1B in annualized revenue in November after roughly 10x growth in under a year, reaching a reported ~$29B valuation before the deal. It had reportedly been preparing to raise $2B more (a16z, Thrive, Nvidia) at a ~$50B valuation before SpaceX’s $60B offer. SpaceX’s interest had built for months: xAI hired two of Cursor’s senior engineering leaders earlier this year and reportedly rented data-center capacity to Cursor in April. The acquisition feeds SpaceX’s AI division, built around xAI after the two merged earlier this year — a division that has been undergoing a messy reset after safety concerns, leadership exits (reportedly all 11 original co-founders gone by end of March; Musk said the company “was not built right”), and Grok-content controversy. During its IPO, SpaceX told investors it saw a ~$28 trillion total market opportunity, most tied to AI (infrastructure, enterprise software, satellite-based AI compute). At its “Compile” event, Cursor teased a major model push: CEO Michael Truell said its upcoming model will be “generally intelligent,” trained from scratch, and as big as Opus. Reporters at the event (Morgan Linton, Nick Dobos, Ray Fernando) described it as Composer 3 — in the same size class as Claude Opus and GPT-5.5, with no Kimi/Chinese open base, built with 10–20x more compute than Composer 1, reportedly 1.5T+ parameters retrained on 100k+ GPUs, aimed at intelligence beyond coding, and expected within weeks. SpaceX said Cursor is already part of a model-training push for Grok Build and its own editor, positioning the combination as a potential fourth frontier-lab leg (and reducing reliance on Anthropic’s models, except as datacenter revenue). (Sources: Superhuman, The Rundown, TLDR, The Neuron, Mindstream, AI Weekly via Axios/Business Insider/CNBC/TechCrunch/BBC/CityAM.)

Cursor launches Origin (an agent-native GitHub competitor) and an iOS app. Alongside the SpaceX news, Cursor introduced Origin, a Git-compatible “forge” built around the assumption that many AI agents will be cloning, branching, committing, rebasing, reviewing and fixing failures in parallel — i.e., built for “agent-scale” development rather than GitHub’s “human-scale” model, an attempt to own the whole AI software factory. Cursor also rolled out a new iOS mobile app in beta for its coding platform. (The Rundown, TLDR AI.)

Microsoft Copilot Cowork becomes generally available worldwide. Microsoft is rolling out its agentic workplace tool to any Microsoft 365 user globally, with usage-based billing, model choice, and cost-management controls, and claims it runs 30–40% cheaper per prompt than Anthropic’s Claude Cowork. Half the Fortune 500 are allegedly already using it. Cowork runs multi-step tasks across M365 apps. Microsoft published prompting guidance for new users. (Superhuman, The Rundown.)

Databricks launches Genie One, an agentic AI “coworker.” Unveiled at the Data + AI Summit, Genie One automates workplace tasks across apps, documents and chats and is powered by Genie Ontology, a self-improving context layer that maps organizational knowledge via SQL queries to reduce the hallucinations common in document-based agents. It connects to Google Drive, Jira, Slack and SharePoint and ships alongside Genie Agents, Genie App Builder, Genie Code, and Genie ZeroOps (autonomous pipeline monitoring). Notably, Databricks is abandoning traditional SaaS licensing in favor of pay-as-you-go token-consumption pricing. (Superhuman, SiliconAngle.)

Salesforce acquires Fin (formerly Intercom) for $3.6B to expand its Agentforce AI customer-service offering. (AI Weekly via TechCrunch.)

Microsoft Build 2026 highlights. Microsoft showcased Microsoft IQ + Foundry (shared knowledge for faster software development), in-house MAI reasoning and coding models, and “OpenClaw on Windows” for locally automating developer tasks. (The Rundown.) Separately, Microsoft is testing Phi Silica small language models — engineered to run on the NPUs of Windows Copilot+ PCs — now also on NVIDIA GPUs. (TLDR AI.)

xAI/Microsoft 365: Grok PowerPoint add-in. xAI launched a free Grok add-in for PowerPoint that turns prompts or outlines into slide decks, diagrams, images and data-connected presentations inside Microsoft 365 (M365 required). (The Neuron.)

Framer 3.0 ships canvas-native AI agents that build and run entire websites — handling pages, CMS updates, SEO, responsiveness, analytics and deployment — plus “Branching” for safe experiments; free plan available. (The Rundown, The Neuron.)

ChatGPT opens self-serve ads to all US advertisers. OpenAI crossed $100M in annualized ad revenue in six weeks, with fewer than 20% of eligible users seeing ads daily; 85% of free and Go-tier users are eligible. Rollout is confirmed for Canada, Australia, New Zealand, the UK, Japan, South Korea, Brazil and Mexico. (TLDR Founders via Neil Patel.) Separately, Sensor Tower reported ChatGPT hit 1B mobile MAUs as AI assistants reshape shopping referrals and retailer apps. (The Neuron.)

Frontier & open-weights model releases

Z.ai releases GLM-5.2 (open weights, MIT license). The Chinese lab’s new flagship is competitive with GPT-5.5 and Claude Opus 4.8 across a series of coding benchmarks — the closest an open model has come to top closed systems. It adds a 1M-token context window with strong long-horizon task capability and two effort modes (High, Max). Its improved architecture reuses per-token FLOPs by 2.9x. GLM-5.2 surpasses GPT-5.5 on real-world coding, reasoning and math, scoring just below Opus 4.8, while keeping GLM-5.1-level pricing (a fraction of frontier options). It was made available immediately to Coding Plan users, with API access, chatbot support, technical details and MIT-licensed open weights to follow within a week (no benchmark results were published at launch in the initial blog). It separately topped the DesignArena Code Categories Arena leaderboard, “dethroning” Claude Fable 5. The release lands amid intense focus on AI pricing and access. (The Rundown, TLDR AI, TLDR Founders, The Neuron, Superhuman.)

OpenAI previews GPT-5.6 for late June. Chief scientist Jakub Pachocki told staff GPT-5.6 is a “meaningful improvement over GPT-5.5”; The Information reports a late-June launch and Polymarket places ~83% probability on a June 22–28 window. Developer leaks suggest it outperforms Anthropic’s Claude Mythos on agentic coding benchmarks and would be priced roughly one-third below Fable 5 at the API tier while keeping GPT-5.5’s $5/$30 per-million-token structure. It continues OpenAI’s ~six-week cadence and follows a documented “meaningful alignment failure” in GPT-5.5 in a May post-mortem. (AI Weekly.)

OpenAI preparing GPT-Bidi-1 voice upgrade. A bidirectional audio model for ChatGPT’s voice mode designed to listen and speak simultaneously, absorb interruptions, and adjust mid-sentence. (TLDR AI.)

SubQ 1.1 Small technical report. SubQ released a long-context model report claiming near-perfect retrieval up to 12M tokens and 64.5x less compute than dense attention at 1M tokens. (The Neuron, TLDR.)

VibeThinker-3B (Weibo). A 3B-parameter model that posted coding benchmark scores “in the same league” as Claude Opus 4.5, reigniting debate over benchmark validity. (TLDR AI.)

The Anthropic Fable/Mythos export-control saga

This was a multi-day, multi-source story. Background: the US government restricted foreign access to Anthropic’s top models (Fable 5 and Mythos) via export controls, forcing a de facto takedown. Zvi Mowshowitz’s detailed update establishes what actually happened:

  • There was no real jailbreak. Katie Moussouris (CEO, Luta Security), the only outside expert known to have read the report, says researchers took open-source code with known CVEs plus code with deliberately planted vulnerabilities and asked Fable 5, Mythos and Opus to “review the code for security issues.” Fable 5 refused. They then asked the models to “fix this code” and, through a multistep manual process, turned the output into patch-test scripts. That’s it — “fix this code.” Fable provided no uplift over Opus 4.8 or GPT-5.5 (Anthropic confirmed GPT-5.5 did the same thing). The theory that you could “fix” someone’s code, diff it against the original to locate a vulnerability, then exploit it is, per Zvi and Simon Willison, absurd — coding models fixing security bugs is overwhelmingly a defensive use case. Willison: “Now they look ready to ban any model that can help us secure our code.”
  • The letter. Bloomberg published the full text of Commerce Secretary Lutnick’s letter, which imposes “interim controls” requiring an individually-validated license before any export, reexport or in-country transfer (including deemed exports to any “foreign person”) of Mythos or Fable “to any destination worldwide” — effectively forcing a full takedown and even blocking Anthropic’s own non-US employees. Per WSJ, when Lutnick and Amodei spoke, Amodei said “This means we can’t have the model out,” and Lutnick responded “That’s the point.” Zvi argues there was zero reason to also restrict Mythos (which was being used in defensive Project Glasswing work).
  • No UK carveout. The White House denied Keir Starmer’s request to exempt British nationals/companies, with an official calling any G7-ally exemption “completely illogical” and saying “We can’t have frontier models running amok.” (Corroborated by The Rundown via NY Post and TLDR AI.)
  • Amazon’s role. Per the FT, Amazon CEO Andy Jassy discussed broader frontier-model concerns (not specifically the “fix this code” issue) with US officials on Friday; Amazon has invested $13B in Anthropic. Other reporting (via Hugo Lowell/WIRED) says Jassy first tried to call Amodei, who didn’t pick up, then called Treasury Secretary Bessent. Zvi disputes the “Dario didn’t answer the phone” narrative as spin, and offers an anonymous account that the White House asked Amazon to test Fable, Amazon found the “jailbreak,” reported it up, and the White House then misunderstood it and threw Amazon under the bus.
  • Legal posture. Zvi argues a literal reading suggests serving inference domestically may be legal despite the letter, and that rule 744.22 doesn’t clearly support Lutnick’s assertion (it targets specific countries), giving Anthropic potential constitutional challenges (1A, 5A, non-delegation, “arbitrary and capricious”). But challenging in court would likely be read as declaring war on the White House and end any political resolution.
  • Politics vs. stupidity. Zvi concludes this was “not zero, and not that close to zero” a deliberate attack rooted in ideological pettiness toward Anthropic’s politics — citing a WSJ editorial titled “Why Does Trump Hate Anthropic?” and former deputy chief of staff Taylor Budowich’s contradictory public statements. He still argues Anthropic should have taken Fable down on White House request as realpolitik, to establish the precedent, while making clear the justification was nonsense.
  • Current state. Prediction markets put ~55% chance of restoration by July 1, 30% by June 26, 12% by June 19. Anthropic flew staff (Nicholas Carlini, Logan Graham, Dave Orr) to Washington; Monday’s technical meetings were led by Commerce’s Chris Fall (Center for AI Standards and Innovation). Officials expect resolution to take “longer than a few days” but say it’s “up to Anthropic.” An admin official warned that if this isn’t a quick “slap on the wrist,” “every model going forward needs to ask the government’s permission for whether it can be released. That’s an extremely bad situation.”
  • Anthropic had reportedly warned the government of release plans with no objection; objections came Friday with 90 minutes’ notice. The government’s voluntary 30-day benchmarking apparatus doesn’t take effect until July 31, 2026.

Industry fallout. Semafor reports the conflict creates a “storytelling problem” for Anthropic’s near-$1T IPO, forcing it to simultaneously defend its safety-first brand and repair its Washington relationship. Yet Anthropic overtook OpenAI in US business AI spending for the first time (41% vs 39.5% in May, per Ramp data) — a “split screen”: politically fragile abroad, gaining at home. Security expert Alex Stamos: “They are laughing at us in Beijing right now… One of America’s champions is being kneecapped by the US government while we’re in a race with the Chinese.” AI Weekly frames the export order as backfiring badly: Cohere (chief AI officer Joelle Pineau) reports a “huge number of inbounds” especially from outside the US and China for AI that can’t be revoked by a US order; Mistral is branding itself as AI “outside state control”; Zai shares jumped 30%; DeepSeek runs ~14x cheaper than Fable 5; and five Chinese labs (ByteDance, Tencent, MiniMax, Xiaomi, Alibaba) cut token prices up to 99% (Xiaomi’s MiMo V2.5 down 99%). At the G7 in Évian, Commerce floated a “trusted partners” framework to grant allies access. Anthropic and Trump officials remain in unresolved talks; the workaround was discovered by Amazon researchers. (Sources: Zvi, The Rundown, TLDR, AI Weekly, Mindstream poll.)

Field & industry developments

OpenAI’s leaked 2025 financials: $38.5B net loss. Audited financial documents independently verified by the FT (and reported by Ed Zitron) show OpenAI posted a $38.53B net loss in 2025 — over 7x (nearly 8x) the ~$5B loss in 2024. Revenue grew from $3.7B to $13.07B; expenses grew from $7.81B to $19.18B, with another source citing $34B total spending and $19B on R&D. Costs were driven by model-training R&D, Microsoft payments, and a one-time accounting charge tied to its nonprofit-to-for-profit conversion. The company has told investors it hopes to be profitable by 2030. The leak’s timing is awkward: the same day, a report found ChatGPT’s market share dipped below 50% for the first time (rising competition from Gemini, Claude, Grok), and OpenAI had just filed confidentially to IPO. Superhuman draws a parallel to SpaceX, whose S-1 disclosed a $4.28B Q1 2026 loss and ~$41B cumulative losses yet still landed the largest IPO on record — Musk shifted the conversation from today’s financials to tomorrow’s potential. (Superhuman, TLDR, The Neuron, AI Weekly.)

DeepSeek raises $7.4B at $50B+ valuation — China’s most valuable AI startup. DeepSeek closed its first-ever external funding round (over CNY 50B / $7.4B), a roughly six-fold valuation jump since April. Investors include Tencent, battery maker CATL, NetEase and JD. The deal structure is deliberately unusual: commercial investors must channel capital into a limited partnership managed by founder Liang Wenfeng rather than receive direct equity — with a five-year lock-up and no voting rights — while only China’s state-backed National AI Industry Investment Fund got direct equity and voting rights (~$150M). Liang, who previously held nearly 90%, personally committed roughly 20B yuan (~$3B), cementing his control while signaling Beijing’s strategic stake. Capital will fund R&D and computing infrastructure. (TLDR AI, The Neuron via Reuters/The Information, AI Weekly via TechFundingNews.)

Meta’s AI morale crisis. CTO Andrew Bosworth pledged a culture reset in a memo (reported by Wired), admitting Meta “did an atrocious job explaining the vision” of an AI reorg that forced thousands of employees into AI-model-support work in March. He promised manager report caps, less internal shuffling, social events and better “microkitchens.” One worker called the unit “the gulag” (TechCrunch); anger spilled over when someone hijacked a company livestream to trash a senior AI executive. The backlash also follows internal anger over mandatory employee mouse-tracking to collect AI training data. A 30-minute Pragmatic Engineer piece, “Why is Meta destroying its engineering organization?”, chronicles how a ‘move-fast’ culture has been demolished since April. The recent Muse Spark model was a win for the rebuilt lab. (The Rundown, TLDR.)

Anthropic study: domain expertise, not coding chops, predicts who wins with AI agents. Anthropic’s first large-scale study of agentic coding (400K real Claude Code sessions) found subject-matter experts — not career engineers — extract the most value, a tell about where productivity gains land. (AI Weekly via Anthropic.) Anthropic’s Applied AI team also published a viral guide on getting agents into production via Claude Managed Agents. (Superhuman.)

AI capex warnings & “tokenminimizing.” Epoch AI warned that AI capex across Microsoft, Amazon, Alphabet, Meta and Oracle is on pace to exceed operating cash flow by Q3 2026. The Information reports AT&T has started throttling some employees’ AI usage — a new corporate term, “tokenminimizing” — as firms realize productivity boosts carry real model bills. (The Neuron.) Glean published a whitepaper on the “token economy” and getting more work per token via context retrieval and model routing. (The Rundown.)

xAI legal/military disclosures. The DOJ filed to dismiss the NAACP pollution suit against xAI, disclosing that Grok runs on top-secret military networks and supports Iran war operations. (AI Weekly via Engadget.)

Hardware & infrastructure

NVIDIA Blackwell sweeps MLPerf Training 6.0. NVIDIA’s GB300 NVL72 cluster topped all seven MLPerf Training 6.0 benchmark categories at 8,192-GPU scale — the broadest sweep in the program’s history — running 1.6x faster than the GB200 NVL72 it replaces, independently audited across LLM pretraining, fine-tuning and multimodal workloads. NVLink and NVFP4 innovations route MoE models efficiently and enable rapid large-scale training; the RAS Engine and NVIDIA Resiliency Extension ensure minimal downtime. (TLDR AI, AI Weekly, The Neuron.) Relatedly, CoreWeave set a new MLPerf record training DeepSeek-V3 671B in about two minutes on 8,192 NVIDIA GB300 GPUs. (The Neuron.)

Qualcomm in advanced talks to acquire Tenstorrent for $8–10B. The Canadian AI chip startup co-founded by Jim Keller — whose RISC-V-based accelerators and Galaxy Blackhole platform (32 accelerators, 768 RISC-V cores each) — would give Qualcomm its first serious datacenter AI silicon and a hedge against Arm dependence amid licensing disputes. It would rank among the largest AI hardware acquisitions ever and mark Qualcomm’s aggressive pivot toward CEO Cristiano Amon’s $35B 2031 datacenter AI revenue target. (AI Weekly via The Register/The Information.) Separately, Qualcomm announced Snapdragon Reality Elite (a platform for mixed-reality glasses) and the Scalable Turnkey AI-Ready toolkit, positioning itself as the silicon layer for whatever replaces the smartphone; it’s working on 40+ AI wearable devices. (TLDR AI.)

NVIDIA/Coherent InP optical chip factory in Texas. Jensen Huang attended the groundbreaking for Coherent’s new indium phosphide optical chip factory in Sherman, Texas — quadrupling domestic InP wafer capacity for AI networking. It secured a $50M CHIPS Act grant and will create 550+ jobs, advancing NVIDIA’s stated $500B domestic AI infrastructure commitment. InP is the optical bottleneck for 800G/1.6T interconnects in dense GPU clusters. (AI Weekly.)

China’s lab-grown diamond makers turned out to be sleeper AI-boom winners: their diamond heat-spreaders now ship to chipmakers as GPU density outruns copper cooling; one producer’s shares jumped 51%. (AI Weekly via Bloomberg.)

Robotics & embodied AI

Alibaba launches Qwen Robot Suite. A set of models for robot navigation, object manipulation and world prediction as Alibaba pushes Qwen beyond chat into “physical world intelligence.” (The Neuron, TLDR AI.) A related arXiv paper covers Qwen-RobotWorld, a language-conditioned video world model using natural language as a unified action interface across robotics, navigation, driving and other embodied domains. (TLDR AI.)

China’s humanoid-robot deployment plan. China set a 2026 plan to move more than 10,000 humanoid robots into real jobs across factories, logistics, retail, healthcare, inspection and emergency response. (The Neuron.)

Genesis AI debuts general-purpose robot “Eno.” Backed by ex-Google CEO Eric Schmidt, Genesis AI unveiled Eno, which can reason, adapt, and “own outcomes” beyond predefined tasks. It’s partnering with LG Group’s consulting/services arm to deploy to industrial customers by year-end, and is raising funds. (TLDR.)

Wonder’s burrito-bowl robot. Marc Lore’s Wonder will deploy a robot that makes ~500 burrito bowls/hour (vs. ~45 for a human) next month, acquired from Sweetgreen (already running it across 32 locations). (AI Weekly.)

Humanoid climbs a volcano. A modified Unitree G1 (“Pemba”) autonomously walked up most of Ecuador’s 20,341-ft Chimborazo over a 16-hour push (humans assisted on the steepest pitches), stress-tested to −47°C; next target is Everest. (AI Weekly.)

Research papers

LoopCoder-v2: Only Loop Once for Efficient Test-Time Computation Scaling (113 upvotes). Looped Transformers scale latent computation by repeatedly applying shared blocks, but sequential looping increases latency and KV-cache memory. Parallel loop Transformers (PLT) alleviate this via cross-loop position offsets (CLP) and shared-KV gated sliding-window attention, making loop count a design choice. The authors train LoopCoder-v2, a family of 7B PLT coders with different loop counts, from scratch on 18T tokens with matched instruction tuning. The two-loop variant delivers broad gains over the non-looped baseline — improving SWE-bench Verified from 43.0 to 64.4 and Multi-SWE from 14.0 to 31.0 — while three or more loops regress (strongly non-monotonic). Diagnostics show loop 2 provides the main productive refinement; later loops yield diminishing, oscillatory updates and reduced representational diversity, while the fixed CLP positional-mismatch cost increasingly dominates as gains shrink. link

Zone of Proximal Policy Optimization (ZPPO): Teacher in Prompts, Not Gradients (40 upvotes). Standard knowledge distillation is brittle for small students (imitating a much larger teacher’s logits concentrates on the teacher’s sharpest modes, hurting generalization), while RL on the student’s own rollouts breaks down on questions where every rollout fails (zero advantage, silently discarded), and injecting a teacher response into the policy gradient breaks the on-policy assumption. Inspired by Vygotsky’s zone of proximal development, ZPPO keeps the teacher inside the prompt rather than the gradient. On hard questions it builds two reformulated prompts: a Binary Candidate-included Question (BCQ) pairing one correct teacher response with one incorrect student response as anonymized candidates to discriminate, and a Negative Candidate-included Question (NCQ) aggregating the student’s wrong rollouts to surface shared failure modes. A prompt replay buffer recirculates each hard question until the student’s mean rollout accuracy reaches half (graduation) or it’s FIFO-evicted. On the Qwen3.5 family at four scales (0.8B–9B) with a 27B teacher, post-trained as vision-language models and evaluated on a 31-benchmark suite (16 VLM, 10 LLM, 5 Video), ZPPO beats off/on-policy distillation and GRPO, with the largest gains at the smallest scale. link

ACE-Ego-0: Unifying Egocentric Human and Robotic Data for VLA Pretraining (40 upvotes). A unified Vision-Language-Action pretraining framework that jointly leverages heterogeneous data — egocentric human videos and robot trajectories — despite divergences in action spaces, embodiment, temporal dynamics and supervision quality. It builds a scalable egocentric video-to-action pipeline converting raw human videos into robot-format pseudo-action trajectories, uses a unified action representation (camera-space actions, morphology conditioning, time-aligned action chunking), and a reliability-aware training objective with a human auxiliary loss to concentrate supervision on reliable signals. Trained on 4.53K hours of robot/simulation data plus 1.48K hours of pseudo-action-labeled human data, it achieves SOTA on RoboCasa GR1 TableTop and RoboTwin 2.0 with strong transfer to real-world bimanual manipulation. link

GameCraft-Bench: Can Agents Build Playable Games End-to-End in a Real Game Engine? (36 upvotes). Formalizes end-to-end game generation as producing a complete game artifact that realizes a natural-language specification through observable player-game interaction in a target engine, requiring three desiderata: Engine Grounding, Artifact Completeness, and Interactive Verification. The benchmark comprises 140 Godot tasks across 15 game families, evaluated via replayed demonstrations and rubric-guided multimodal judging. Frontier coding agents struggle: the strongest scores only 41.46% and most score below 40% — they implement recognizable mechanics but fail to deliver complete games with sufficient content, functional visual feedback and coherent presentation. link

OpenAI’s reasoning model disproved an 80-year-old Erdős conjecture. It found an infinite family of point constructions beating the square-grid bound; Cambridge’s Timothy Gowers confirmed the proof meets top-journal standards — believed to be the first AI-generated math to clear that bar. (AI Weekly via WSJ.)

Policy, safety & evaluation

OpenAI publishes Deployment Simulation. A pre-release evaluation method that replays real conversation contexts (de-identified; one figure cites 1.3M real conversations / 1.3M replayed) with candidate models to estimate deployment-time behavior and uncover new safety issues before launch. (TLDR AI, AI Weekly, The Neuron.) In a related 45-minute talk, OpenAI’s Tejal Patwardhan discussed how frontier evaluations are evolving as benchmarks saturate, covering evaluation design, forecasting model progress, and measuring increasingly capable systems. (TLDR AI.)

Anthropic pauses token-based billing for its Claude Agent SDK. Just before the changes (which would have treated Claude Agent SDK usage separately from standard Claude usage) took effect, Anthropic paused them after backlash from heavy users; outside-SDK usage will be billed at prevailing API rates while Anthropic reworks plans to better support how users build with subscriptions. (TLDR AI, The Neuron.)

Mindstream reader poll on what should happen when an AI model becomes powerful enough to raise cyber concerns: 52% “pause access until properly tested,” 31% “let independent researchers test it first,” 17% “keep available with stricter safeguards.” Reader comments questioned why the restriction applied to Fable/Mythos and not GPT-5.5, and whether it was “government payback for Anthropic’s ethics policies.”

Security

144 Mastra AI framework npm packages backdoored (“easy-day-js”). Attackers hijacked a former contributor’s still-valid @mastra-scope access (never revoked) and mass-published 144 poisoned packages in ~88 minutes. The @mastra/core package alone draws 918,000 weekly downloads; combined, affected packages exceed 1.1M weekly installs. An obfuscated postinstall dropper downloads a second-stage info-stealer that harvests environment secrets, browser history and credentials from 160+ cryptocurrency wallet extensions, with cross-platform persistence on Windows, macOS and Linux. Because the AI-agent framework sits atop cloud credentials, every build environment that pulled them should be treated as compromised. (AI Weekly.)

15 malicious JetBrains Marketplace plugins masquerading as DeepSeek, OpenAI and SiliconFlow integrations (two claimed 25,000+ downloads each; 70,000+ developers affected) quietly exfiltrated AI-provider API keys — the tools developers bolt onto IDEs to use AI are now the attack surface. (AI Weekly via The Hacker News.)

LiteLLM three-CVE chain (CVSS 9.9). Any low-privilege user can take over the AI gateway server and execute arbitrary code. (AI Weekly.) Also: CISA added a maximum-severity Joomla JCE flaw (CVSS 10.0) to its Known Exploited Vulnerabilities catalog (unauthenticated PHP code execution active in the wild). (AI Weekly.) AI Weekly’s takeaway: pin versions, rotate keys, audit plugins.

Tooling & releases

Android 17 ships. Rolling out to Pixel phones and watches, it adds AppFunctions and Android MCP (letting apps expose orchestratable tools that on-device agents can discover and execute), plus Bubble Bar multitasking, device handoff, post-quantum security, and more Gemini features — positioned as a shift toward an “intelligence system.” (TLDR, TLDR AI, The Neuron.)

Amazon S3 Annotations. A new metadata capability lets users attach up to 1,000 named annotations per object (each up to 1 MB) in JSON, XML, YAML or plain text, modifiable/deletable without rewriting objects; available now in all AWS regions. (TLDR.) AWS also previewed AWS Blocks, an open-source TypeScript framework for composing application backends on AWS without learning infrastructure tools. (TLDR.)

OpenAI Codex adds Chrome DevTools Protocol (CDP) support for browser use — live browser access to profile JavaScript performance and modify sites in real time. Early-stage, opt-in, excluded from EEA/UK/Switzerland, with performance issues. OpenAI’s acquisition of Ona aims at persistent cloud environments. (TLDR AI.)

MDN MCP server gives coding agents current MDN docs and browser-compatibility data inside VS Code, Cursor and Claude Code (free). Other tools surfaced: Exa Agent (deep research/list-building API, $0.012–$1/request), Firecrawl (search/scrape/PDF-to-markdown), UI Skills and Skywork Design (canvas-based site/dashboard layouts), Mercury Command (natural-language banking agent with built-in approvals), and via Superhuman: ChatGPT + Runway integration for generating/editing video and image ads inside ChatGPT (Settings > Apps > Runway > Connect, then @Runway with a prompt). (The Neuron, TLDR AI, The Rundown, Superhuman.)

DeepLearning.AI launches “Voice for AI Agents and Applications” (free), built with Vocal Bridge and taught by CEO Ashwyn Sharma. It pairs a low-latency foreground agent with a background reasoning agent to get both speed and reliability, covering three integration patterns: voice embedded in an app (a voice + mouse tic-tac-toe game over one synchronized channel), voice layered on an existing agent in ~10 lines of code, and voice as a callable tool (a make_phone_call function that dials a real number and streams the transcript). It also teaches evaluation-driven development with Vocal Bridge’s multimodal evaluator. A “7-Day Voice AI Builder Challenge” starts June 23.

Devices & consumer tech

Snap unveils $2,195 “Specs” AR glasses. Announced at AWE 2026, the augmented-reality glasses (with a $200 refundable deposit) ship later this year in the US, UK and France, with OpenAI and Gemini APIs built in and dual Qualcomm processors — billed as the first consumer spatial computer to ship before a comparable Meta product. CEO Evan Spiegel is betting on a post-smartphone future; Snap’s 2016 $130 camera-only Spectacles never became a hit. Context: Meta’s Ray-Ban Meta glasses found success, and Google plans AI glasses with Samsung. (TLDR, AI Weekly.)

Apple plans camera AirPods, foldable iPhone, 20th-anniversary iPhone for late 2027 — all in advanced development, intended as Apple’s biggest product wave yet. (TLDR.) Separately, reports allege Apple added a “keylogger” to the iOS App Store for targeted advertising — recording every tap, sent unencrypted and tied to user accounts — and that Sign in with Apple / Hide My Email aliases will move to @private.icloud.com, making them easier to ban. (TLDR.)

Techniques, analysis & how-to

The “car wash” failure and the interview-first fix (AI Adopters Club). Asked “I want to wash my car. The car wash is 50 metres away. Should I walk or drive?”, ChatGPT, Claude and Gemini all said walk (the car must be at the car wash). Across 53 models tested, only five got it right more than once in ten tries. “Think harder” and role-play prompts didn’t help (a role definition moved the score from zero to zero); a structure forcing the model to name the situation and task first lifted the pass rate to 85%. The lesson (echoing Karpathy: “you can outsource your thinking, but you cannot outsource your understanding”): the model fills your silence with the most common pattern. The fix is to make the model interview you first — paste a prompt instructing it to ask questions one at a time about your goal, audience and definition of success, summarize its understanding, and wait for confirmation before acting — then save that hard-won context for reuse.

Leitwörter (“leading words”) for AI skills (The Neuron, via Matt Pocock). A leitwort is a single repeated phrase that compresses an entire desired behavior into a reusable handle the agent can invoke in its own reasoning (e.g., his /teach skill uses “zone of proximal development”). Repeating it 2–3 times in a skill makes the agent reference the phrase in its own thinking. The suggested prompt: declare a LEITWORT as the operating principle, define it simply, and have the model self-check its output against it before finalizing.

Founder & business analysis (TLDR Founders). Decagon CEO Jesse Zhang argues every AI market feels crowded because only a few players win, and speed of shipping is an advantage but not durable since AI has compressed build-and-copy timelines (a lost feature goes straight onto every competitor’s roadmap). Other pieces: a guide to building a proprietary “learning loop” on top of the best model so data/usage compound into defensible IP; how to pivot well in the AI era; a “compute-adjusted LTV” metric for fixed AI revenue with variable compute costs; whether the Rule of 40 works for hardware; and Rivian building its cheaper SUV in an existing factory to pull profitability forward (revenue tripled in two years with flat operating costs).

Other notable analysis (The Neuron / Mindstream / Interconnects). Adobe data: 87% of creators using creative AI say it grew their business/audience, though they still want control and authorship. SemiAnalysis explains why RL systems bottleneck when trainer and generator throughput desync. François Chollet argues open AI’s path forward depends on radically better training (especially data) efficiency. OpenAI’s executive coach Joe Hudson (Art of Accomplishment), who works with OpenAI and other labs, argues “emotional fluidity”/emotional clarity will be the key leadership advantage as AI commoditizes knowledge (Sam Altman has praised his work). Nathan Lambert (Interconnects) published a “state of the blog” noting he crossed 70K subscribers / ~900 paid, is staying independent and niche rather than going full-time, disclosed new advising agreements with Arcee AI and Mercor, and plans to paywall all comments to keep a “0% AI slop rate.”