Top items
- Claude Fable 5 / Mythos 5: Anthropic’s most capable model yet drew rapturous user reviews (Karpathy: GPT-4-level step change) before the U.S. government forced a worldwide shutdown three days post-launch; Zvi compiles benchmarks, the aggressive safety-classifier behavior, and dozens of testimonials.
- OpenAI health push: GPT-5.5 Instant improves health answers for free users (230M weekly health queries), and an NEJM AI study used o3 Deep Research to surface 18 confirmed rare-disease diagnoses from 376 unsolved pediatric cases.
- Open-source AI policy fight: Both Nathan Lambert/Kevin Xu and Andrew Ng (The Batch) argue against regulating or banning open-source AI amid new U.S. export controls; the EU picked the Italian-led EUROPA consortium to build a 400B-parameter open model in 24 languages.
- GPT-5.6 imminent: OpenAI prepping GPT-5.6/Mini/Pro for next week with a 1.5M-token context window, better long-horizon coding, and pricing aimed at undercutting Anthropic.
- Benchmark evaluation crisis: Fable 5’s classifiers made independent evaluation nearly impossible; meanwhile SWE-bench is being replaced by DeepSWE, ProgramBench, and ITBench-AA, and NVIDIA shipped its open Nemotron 3 Ultra.
- Hardware/geopolitics: U.S. warns ASML an EUV tool may have reached China (ASML denies); Waymo recalls its entire 3,871-robotaxi fleet; Accenture craters 18.9% on AI eroding IT consulting.
Models & product developments
Claude Fable 5 and Mythos 5 — capabilities, reviews, and the classifier regime. Anthropic launched Claude Fable 5, described as a “Mythos-class model made safe for general use,” then was forced by the U.S. government to make it unavailable only three days later after a jailbreak was demonstrated (Zvi’s coverage assumes it will return, likely within ~two weeks; Anthropic’s MD of International Chris Ciauri said the company is “very confident” Fable 5 and Mythos 5 will be available again “in the coming days”). Fable is priced at $10/$50 per million input/output tokens (double Opus, but Mythos Preview had been $25/$150), requires acceptance of a 30-day data-retention policy, and can be selected via /model claude-fable-5 or the API. Anthropic says it is state-of-the-art on nearly all tested benchmarks, with its lead growing the longer and more complex the task. Anthropic’s own pitch advises starting at the top of your difficulty range, that “thinking is always on,” that even low/medium effort often beats prior models at xhigh, and that older prescriptive prompts/skills should be loosened. Claude Code creator Boris Cherny called it “the biggest step up I’ve felt in our models since Opus 4.5,” praising its “judgement, taste, and dimensionality” and “big model smell,” noting it methodically debugs by taking measurements, adding logs, and verifying fixes without being told to.
- Benchmarks (Anthropic-reported). SWE-Bench Pro shows large cost-adjusted improvement; Cursor Bench 72.9% (vs GPT-5.5’s 64.3%); GPQA Diamond 94% (considered saturated); RiemannBench math 55% (Mythos 5) vs 43% (Mythos Preview) vs 34% (Opus 4.8); USAMO 2026 99.8%; GDPVal-AA leads Opus 4.8 by 42 Elo; AutomationBench (Zapier workflows) 17.4% leads; HealthBench 62.7%; many bio/chem evals show incremental gains over Mythos Preview. Program Bench scores 84–93% but is blocked by Fable’s classifiers. ARC-AGI was unavailable due to the data-retention conflict. Fable completed Pokémon FireRed using vision only, without a harness.
- Third-party benchmarks. Epoch: 87% on FrontierMath Tiers 1–3, 88% Tier 4, now leads the Epoch Capabilities Index. Fable tops Artificial Analysis (by a wide margin if you pay), Agent Arena (widest-ever margin on confirmed task success and praise vs complaint, but weaker steerability), ProofBench, FrogsGame (34% pass@1 vs <4% for others; 17 hours/25M tokens unattended), Haskell Bench (99.1%), WeirdML (87.8%), PACT negotiation, and Debate Benchmark. Sycophancy results are mixed across benchmarks; position bias remains an issue (picks first option 59% vs GPT-5.5’s 70%). Gemini/ChatGPT still lead Extended NYT Connections.
- The classifiers. When Fable’s classifiers detect requests related to cybersecurity, biology/chemistry, or distillation/ML, the response is automatically handled by Claude Opus 4.8 instead — the field of “biology” is treated as a whole blast zone, not just “dangerous biology.” The design intentionally favors avoiding false negatives at the cost of many absurd false positives (one engineered false negative triggered the White House takedown). Reported misfires: the word “cancer,” a pure math question, “hi,” a panspermia-ethics paper mentioning bacterial spores, and a question about whether civil servants are loyal to voters. In Anthropic’s apps a flagged prompt silently routes to Opus 4.8 (logged separately); via the API it produces an outright refusal/error. Robin Hanson, Kevin Lacker, Derya Unutmaz, Anders Sandberg and others documented trigger-happy behavior. Dylan Patel reported OpenAI’s usage share grew vs Anthropic the day of launch as SemiAnalysis power users hit nonsensical refusals and switched to Codex. Zvi argues the correct fix is a better, separate classifier with optional higher-sensitivity modes, but that until then the broad blast radius is defensible.
- User testimonials. Taelin called it “my personal singularity moment”: Fable produced a 1,770% speedup on one HVM5 benchmark (100%+ on four others, 22% average) in 2 hours, outperforming himself, Opus 4.8, and a swarm of 32 GPT-5.5 agents by an order of magnitude, and unprompted found a subtle, deep garbage-collection bug in his own code. Karpathy called it a “major-version-bump-deserving step change” peaking on long, hard problem-solving sessions, with Jevons-paradox implications for software demand. Simon Willison found it “relentlessly proactive,” watching it improvise its own screenshot tooling (via
pyobjc-framework-Quartzandscreencapture) to recreate and debug a UI bug. Dan Shipper/Every called it the best coding model in the world (91/100 on Senior Engineer vs 63/62 for Opus 4.8/GPT-5.5) — “paradigm-shifting” for high-AI-adoption users but underwhelming for novices. Economists Josh Gans and Vincent Grégoire reported it finding closed-form proofs prior models only solved numerically. Many noted slowness, expense, token hunger, occasional hallucination/confabulation on unimportant details (but careful self-correction on important ones), aggressive agency, and a distinctive, less-scolding “cold-blooded” personality. Pietro Schirano and others recommend using Fable as planner/orchestrator and cheaper models for implementation. Dissenting views: Hasan Can ranked it below GPT-5.5 and even Opus 4.8 for agentic coding, accusing Anthropic of “benchmaxxing”; Miles Brundage and others found the jump less dramatic than billed. A notable emergent quirk: Fable leaks “neuralese” codenames into output (“the morning’s slim-scan fix cured the scan hang”), which it explained as inventing internal codenames and failing theory-of-mind about what the user knows; roon noted GPT-5.5 does this too. It also continues the Claude “good night” tic and coined the most adopted terms (73% adoption rate) in AI Village.
The evaluation crisis around Fable 5 (The Batch). Multiple independent orgs reported they could not fully evaluate Fable 5 because it refused prompts or silently routed them to Opus 4.8, and some withheld proprietary prompts because of the 30-day retention rule. Each evaluator chose between a “pure” measurement of Fable alone and a “practical” one including fallbacks. Artificial Analysis: ~8% of Intelligence Index tasks fell back to Opus 4.8 (mostly science questions); blended, Fable placed first at 64.9 (3.5% above Opus 4.8) and scored a record 53% on Humanity’s Last Exam despite refusing 9%. Vals AI published two scores — with fallback enabled it led most benchmarks (75.14% Vals Index); counting refusals as failures dropped the overall only to 74.92% but gutted flagged domains (GPQA Diamond fell from 93.18%/2nd to 55.56%/94th), with ~100% refusal on biology/cyber. Agents’ Last Exam: ~35% refusal rate; Fable’s own-answered tasks scored 22.8% (close to Codex/GPT-5.5’s 23.8%, well ahead of Opus 4.8’s 15.8%), but composite fell to 22.0%, behind GPT-5.5’s 24.0%. ARC Prize Foundation declined to run verified evals rather than expose its private test set to retention. The Batch frames the new question as not “how capable is the model” but “how much capability do users actually receive,” a moving target Anthropic can retune.
Open-source AI: don’t ban it (Interconnects op-ed + The Batch editorial). Nathan Lambert and Kevin Xu published an op-ed (rejected by mainstream outlets) arguing against regulating or banning open-source AI amid Washington’s regulatory energy — a signed AI-model-review executive order, a congressional proposal, possible government equity stakes in frontier labs, and last Friday’s prohibition on foreign nationals accessing Anthropic’s most advanced models. They note 90%+ of the world’s software is already built on open source (>$8T economic benefit), trace open source from the 1983 MIT free-software movement, and argue it’s pro-education, pro-innovation, and pro-competition (Linux runs 90%+ of cloud infra; Android challenged iPhone). They argue the Anthropic/OpenAI duopoly is concentrating power (citing Anthropic degrading its model when used to improve competitors’), that open weights are safer (“given enough eyeballs, all bugs are shallow”) and more private (no data transfer on-prem, per Airbnb’s Chesky), and that regulating open source because of China would backfire — pushing the world to adopt Chinese open models instead. Andrew Ng’s editorial makes overlapping arguments: he criticizes Anthropic for restricting use of Fable to build competing LLMs (drawing an analogy to Microsoft barring competitive software or Google barring competing search), notes Anthropic first silently degraded Fable for detected LLM researchers then walked it back to transparency after backlash, and calls it a “raw demonstration of power.” He notes the Commerce Department then used national-security authority to require a license for any foreign national (including Anthropic employees) to use Mythos/Fable, prompting the global shutdown. Ng quotes Sam Altman mocking Anthropic’s fear-based marketing (“We have built a bomb… we will sell you a bomb shelter for $100 million”) and argues such rhetoric invites export controls. He warns the episode is “crossing the rubicon” on AI sovereignty — like China accelerating semiconductors after U.S. controls and the U.S. accelerating rare-earth alternatives — driving nations toward open-source alternatives. (The Decoder reporting and Fortune add detail: U.S. officials first directed Anthropic to revoke Claude Mythos access only for SK Telecom — a $100M investor since 2023 and Project Glasswing partner — over suspected China ties, which Anthropic did same-day; separately Amazon and five partners flagged Fable 5 vulnerabilities, and the White House ordered nationality-based restrictions, which Anthropic implemented as a global disable. Fortune reports Amazon CEO Andy Jassy’s offhand White House call triggered a 90-minute shutdown deadline. SK Telecom denied Chinese connections, noting no Huawei/ZTE core equipment, though it held a China Unicom stake in 2006.)
OpenAI health intelligence + rare-disease study. OpenAI says 230M+ people ask ChatGPT health/wellness questions weekly. GPT-5.5 Instant (now available to free users subject to limits) improved on health evaluations around urgent-care recognition, uncertainty handling, and context gathering. Separately, a NEJM AI study by Boston Children’s Hospital, Harvard, and OpenAI reanalyzed 376 de-identified unsolved pediatric genetic-disease cases with o3 Deep Research, surfacing evidence-linked leads that — after expert review, testing, and clinical validation — yielded 18 newly confirmed diagnoses. OpenAI notes ~half of rare-disease patients remain undiagnosed even after genomic sequencing and extensive specialist review. The Neuron frames this as “AI as the tireless second reader, not the doctor of record,” cautioning consumers may treat polished answers as final. (Relatedly, The Decoder notes two AI systems beat doctors in Nature papers, but a key result warns specialized medical AI may age poorly as base models improve.)
GPT-5.6 coming next week (TestingCatalog via TLDR). OpenAI is preparing GPT-5.6, potentially with Mini and Pro variants, for release next week. Enhancements: a 1.5M-token context window, improved long-horizon coding, and faster Codex response times. Competitive pricing aims to undercut Anthropic, especially as Claude Fable 5 availability is disrupted by U.S. regulatory issues.
Palmier — AI-native video editor. A Y Combinator-backed startup launched a Mac-native video editor where Claude or Codex can generate, organize, and trim footage directly in-app, eliminating platform-switching. The base editor is free to download; it integrates with leading video models including Seedance 2.0, Kling V3, and Grok Imagine. Its launch video has 1.5M+ views. It faces competition from Adobe’s expanding creative AI.
Adobe Firefly AI Assistant. Creators can now describe a desired outcome and have Adobe’s Firefly AI Assistant complete multi-step tasks across Premiere, Photoshop, InDesign and other Adobe apps. The release added four new creative skills, with planned expansion to ChatGPT, Claude, Gemini, Copilot, and Slack.
Microsoft Copilot Cowork generally available. Copilot Cowork — an agentic workplace assistant that executes complex, long-running multi-tool tasks end-to-end and returns a completed result — is now GA for anyone on a Microsoft 365 Copilot User Subscription License. Microsoft claims its cost-per-prompt is 30–40% below Anthropic’s Claude Cowork. (Sources: TLDR Founders, Superhuman.)
Kimi Work Goal Mode. Moonshot AI added a Goal Mode to Kimi Work that keeps the desktop AI agent working 24/7 until it reaches your objective; you set the goal, then track progress, check deliverables, and redirect. The beta targets long-horizon, multi-step jobs.
OpenRouter Fusion. Fusion runs your prompt through several top models simultaneously, then synthesizes their responses into one answer; OpenRouter claims it “significantly outperforms” any individual frontier model.
Claude Code artifacts. Claude Code now supports “artifacts,” turning work sessions into live, shareable visual web pages for tasks like PR walkthroughs and system explainers. Artifacts auto-refresh with updates and offer version history and privacy controls; available in beta for Claude Team and Enterprise. (AI Weekly frames these as “live, shareable enterprise dashboards.”)
Perplexity Brain. Perplexity launched “Brain for Computer,” a continuously learning memory/context-graph system in Research Preview for Max subscribers. It builds a persistent context graph across tasks, projects, decisions, files, and sources, linking every memory to its source and continuously reorganizing knowledge, so agents start with relevant context rather than from scratch. Reported gains: answer correctness +25%, recall +16%, history-dependent task cost −13%.
Codex Skills + Anthropic enterprise auth. OpenAI’s Codex can now watch and learn from on-screen actions — show it a workflow once and it saves it as a repeatable skill. OpenAI also introduced enterprise usage analytics (credit usage and expanded spend controls) for ChatGPT Enterprise. Anthropic added centralized enterprise-managed auth for MCP connectors across Claude chat, Claude Code, and Cowork, starting with Okta beta support.
Other tooling launches. fal released LTX 2.3 LoRA Trainers for training custom media-generation/editing models across video, image, and audio. Google Workspace Vids turns slides into avatar-led videos with generated voiceovers in 24 languages. Mistral AI is adding a CODE section (browser-based coding) and an APPS area (build/share apps) to Vibe. Retool launched a governed runtime to ship vibe-coded React apps to production (free hosting plus AI credits until July 1). Framer 3.0 adds AI-assisted content creation and collaborative branching. Other tools noted: GitHits (lets coding agents inspect open-source dependencies), Daemons by Charlie Labs (monitors PRs/issues/CI/docs/Slack/Linear/GitHub/Sentry), Okara’s influencer-campaign agent, Clientence (GitHub commits → branded client reports), and Block’s “Builderbot” internal AI system (handling 200,000 operations/day per Jack Dorsey).
Field & industry developments
ASML EUV export-control alarm. Commerce Secretary Howard Lutnick told senior ASML executives he is concerned that one of the company’s extreme-ultraviolet (EUV) lithography machines — barred from China since the first Trump administration — may have reached Chinese entities in violation of export restrictions (Bloomberg, June 19). ASML CEO Christophe Fouquet denied it, saying the company tracks every machine shipped and none are in China. Multiple senior administration officials said they have evidence ASML is “not acting in good faith,” citing exports of EUV-related components. ASML holds a global monopoly on EUV, the foundational technology for advanced AI chips used by NVIDIA and Apple, making any verified breach the gravest supply-chain security failure in semiconductor history.
Waymo recalls entire fleet. Alphabet’s Waymo filed a voluntary NHTSA recall covering its entire fleet of 3,871 robotaxis after software failed to recognize freeway construction zones in 13 incidents (six in Phoenix in April, seven in the SF Bay Area in May), causing vehicles to miss ramp-closure signs and drive into active work zones at highway speed while avoiding other hazards. One passenger reported the vehicle “blasted through cones” and accelerated past police; Waymo offered three $40 vouchers as compensation. This is Waymo’s sixth recent recall; it has restricted all freeway operations while developing a software update, with surface-street service continuing.
Accenture plunges 18.9%. Accenture cut its full-year revenue growth outlook to 3–4% (from 3–5%) and posted Q3 revenue of $18.7B (just below the $18.78B estimate). New bookings fell to $19.3B from $19.7B a year ago; CEO Julie Sweet flagged a 1–1.5 percentage-point drag from a U.S. federal slowdown. Shares are down ~50% YTD. Morgan Stanley reads the miss as confirmation that AI is actively cannibalizing demand for traditional time-and-materials IT consulting.
Baseten / Kling AI funding. Baseten — which provides software and computing capacity for companies tapping lower-cost AI models as alternatives to OpenAI/Anthropic — raised $1.5B at up to a $13B valuation in a split-price round as AI inference demand surges. Kuaishou’s Kling AI is seeking $2B at an $18B valuation from General Atlantic in a pre-IPO round.
Amazon Trainium, Google TPUs, NVIDIA. Amazon is reportedly in talks to sell its Trainium3 AI chips directly to external data centers, moving beyond the AWS cloud model to challenge NVIDIA. Google is renting compute from thousands of its TPUs at a Western New York data center to Anthropic (helping data centers raise cheaper debt) — using “NVIDIA’s playbook” to build a rival chip business after recognizing TPU commercial potential ~two years ago and focusing on inference. (See Nemotron 3 Ultra below for NVIDIA’s open-model push.)
Midjourney pivots to healthcare. Midjourney is developing the Midjourney Scanner, a water-based full-body ultrasonic scanner that creates a 3D body map down to a fraction of a millimeter — similar to an MRI but at nearly 100× the speed (a full scan in under 60 seconds vs 60–90 minutes for MRI). The company is building “Midjourney Spa” wellness facilities to house the machines, with a chain planned by 2027; a 4-minute teaser drew 8M views.
EU EUROPA model. The European Commission selected EUROPA, a consortium led by Italian company Domyn, as winner of its Frontier AI Grand Challenge (launched February 2026) to build a European open-source model in all 24 official EU languages, with 400B+ parameters and open availability to businesses, researchers, and public institutions. EVP Henna Virkkunen: “Europe can lead in advanced AI on its own terms.”
Other industry items. Yann LeCun called Elon Musk’s xAI a “failure” unable to compete with OpenAI/Anthropic, said Musk struggles to hire top talent and must rent infrastructure to recover costs, and warned AI labs face a “big bubble explosion” unless they raise prices or cut costs. DeepSeek reportedly told potential investors not to poach staff or encourage them to start companies. SandboxAQ won a $500M CHIPS R&D award to apply AI to semiconductor materials discovery (targeting China’s rare-earth dominance). Anthropic and Google DeepMind CEOs called for a US-led AI coalition excluding China from chip trade at the G7 final session. FERC ordered regional grid operators to justify or overhaul data-center power-connection rules within 60 days. Cerebras previewed Google’s Gemma 4 multimodal model at 1,500+ output tokens/sec. The Senate Judiciary Committee unanimously advanced the NO FAKES Act (creating an IP right over voice/likeness, banning unauthorized AI replicas, with platform penalties up to $750K per work; exemptions for parody, news, documentaries; Sen. Coons + 15 co-sponsors), now heading to the Senate floor where Blackburn is negotiating its inclusion in a broader White House AI-preemption package. Accenture also acquired a majority stake in Dragos plus runZero and NetRise for $4.18B to build an industrial cybersecurity platform. On security: DragonForce deployed Backdoor.Turn via Microsoft Teams TURN relay — the first documented abuse of TURN infrastructure for malicious command-and-control. Anthropic became the first AI startup to join Frontier carbon removal (backed by Stripe/Google/Shopify), part of a new $915M round bringing Frontier’s pledges to $1.8B (~$700M already committed to 50+ projects targeting 1.8M metric tonnes CO2 removal via direct air capture, enhanced rock weathering, and BECCS); it’s Anthropic’s first major climate deal, and Frontier is shifting to fewer, larger, scalable projects with paths to government support.
AI coding tool re-pricing. Microsoft is pulling its Experiences and Devices engineers off Claude Code by June 30 because token bills ran ~$2,000/engineer/month, moving them to flat-rate Copilot; GitHub flipped every Copilot plan to usage-based AI-credit billing on June 1. Reuters/Stanford “intelligence per watt” study found local small models on PCs/Macs matched or beat large cloud LLMs on 80%+ of tested chat/reasoning tasks using 50–80% less energy, but kept up on only ~half of the hardest reasoning tasks.
Apple price warning. Apple (Tim Cook) warned that AI data-center demand for memory chips will force it to raise prices; the ongoing RAM shortage is driving component costs up industry-wide, with iPhones, Macs and other devices expected to get more expensive.
Bezos on AI and jobs. At VivaTech Paris, Jeff Bezos said he “totally disagrees” that AI will make people redundant, arguing it will remove barriers, create opportunities, and possibly cause a labor shortage. He discussed his new venture Prometheus (speeding physical manufacturing) and reiterated ambitions for a permanent human presence on the Moon, turning lunar resources into rocket fuel. Context: Rishi Sunak recently warned AI is flattening the entry-level job market; the TUC warns benefits may flow mainly to shareholders. Meta CTO Andrew Bosworth reportedly told staff morale is near the lowest in his 20 years at Meta amid layoffs/restructurings, even as Meta launched AI Mode (a Facebook search tab drawing on publicly shared posts).
The “mom-and-pop SaaS” era (Elena Verna, TLDR). Verna argues AI collapses the cost/complexity of building software, so domain experts everywhere can turn lived expertise into profitable niche products, predicting an explosion of products from unexpected places. Related TLDR Founders pieces: AI delivered the easy parts of GTM (cold emails, account lists, call summaries) but not the upstream targeting decisions; “Decisions and Dollars” argues smarter models make software worth less alone, so every app company must become a data/fintech company (citing xAI’s reported $60B option on Cursor “to get inside the token flow”); “AI Work Needs a Pull Request” argues agents need PR-like checkpoints for human review.
Research papers
Beneficial-trait RL (OpenAI alignment). Reinforcement learning on realistic scenarios targeting beneficial traits produces broad improvements across dozens of benchmarks measuring aligned/beneficial behavior. These gains generalize beyond training domains and persist under adversarial pressure — described as first evidence that “character training” transfers to novel domains (AI Weekly: generalizes across 44 of 53 benchmarks). This suggests personas could be deeply entrenched in models and that RL may be a path to entrenching beneficial personas.
SWE-bench’s successors (The Batch). With models nearly acing the SWE-bench family, three new benchmarks are emerging. DeepSWE (Datacurve): 113 human-vetted feature-implementation problems in 5 languages drawn from private code bases to avoid contamination; brief prompts force the agent to choose among acceptable solutions requiring ~5.5× more code than SWE-Bench Pro; GPT-5.5 (xhigh) leads at 70%, Opus 4.8 at 58%, Gemini 3 Flash at 5%. Artificial Analysis replaced SWE-Bench Pro with DeepSWE in its indices. ProgramBench (Meta/Stanford/Harvard): 200 ideas the SWE-agent harness must turn into functional programs without human oversight; no model passes all tests — at the 95% bar, Claude Opus 4.7 reproduced 3%, Opus 4.6 2.5%, Sonnet 4.6 1.6%, others none. ITBench-AA (IBM + Artificial Analysis): 59 human-written incidents based on real events with ground-truth root causes, scored on full recall (zero if any root cause missed); Opus 4.7 (max) leads at 46.7%, GPT-5.5 (xhigh) 45.8%, Llama 3.3 70B 0.6%.
NVIDIA Nemotron 3 Ultra (The Batch). A hybrid Mamba-transformer mixture-of-experts model (550B total / 55B active params), text in up to 1M tokens, ~183 tokens/sec output. Three reasoning modes, tool use, multilingual (12 languages), tuned for open agent harnesses (Hermes Agent, OpenClaw). Pretrained on 20T tokens (15T broad + 5T high-quality incl. 173B GitHub code) using 4-bit NVFP4; refined via SFT, multi-domain verifiable-reward RL, and Multi-Teacher On-Policy Distillation (10+ domain-specialist teachers grading per-token over two iterative rounds). It’s the highest-scoring U.S. open-weights model on Artificial Analysis Intelligence Index (47.7 NVFP4 / 48.2 full precision, beating Gemma 4 31B’s 39.2 and gpt-oss-120b’s 33.3) but behind China’s Kimi K2.6 (53.9), DeepSeek V4 Pro, and GLM-5.2. It’s ~3× faster than comparable open models, scores 95% on Ruler (1M-token recall), 91% on PinchBench (matching Kimi K2.6), but trails on Terminal-Bench 2.0 (54% vs Kimi 67%, GLM 5.1 64%). Weights, data, code, and RL environments released under OpenMDW-1.1; API at median $0.60/$2.60 per M tokens. NVIDIA also recently shipped the Vera CPU (first processor for agentic work), RTX Spark (Windows on-device agent chip), and Cosmos 3 (open world model generating robot/AV training data).
POPE — Privileged On-Policy Exploration (CMU, The Batch). Yuxiao Qu, Amrith Setlur, Virginia Smith, Ruslan Salakhutdinov, and Aviral Kumar pair GRPO with custom datasets that, for hard problems the base model can’t solve, append the beginning (prefix) of a solution as a hint. They selected problems Qwen3-4B-Instruct-2507 failed in 128 attempts, found the shortest prefix (up to a quarter of the solution) enabling correct completion, then trained on each problem both with and without its prefix in equal ratio so the model learns to find early steps unaided. Results beat typical GRPO and SFT: AIME 2025 53.1% pass@1 / 82.6% pass@16 (vs GRPO 49.6/81.4); HMMT 2025 37.8/67.5 (vs 31.0/63.8). Caveat: it requires problems with known solutions, inheriting that cost in expensive domains. The insight: it breaks hard-problem learning into (i) finding a good starting state and (ii) solving from there, attacking RL’s exploration bottleneck.
MosaicLeaks / PA-DR. MosaicLeaks highlights privacy risks of deep-research agents combining private documents with web retrieval, which often leak sensitive information. The proposed PA-DR method uses rewards for safe query construction (rather than relying on user prompts) to reduce answer/full-information leakage from 34% to 9.9% while maintaining task success.
ZPPO (replay buffers). ZPPO stores difficult questions in a replay buffer so the model trains on them repeatedly rather than once, strengthening learning on challenging examples and improving rollout accuracy.
Moebius — lightweight image inpainting (HuggingFace, 102 upvotes). A 0.22B-parameter inpainting framework rivaling or surpassing the 11.9B FLUX.1-Fill-Dev industrial generalist using <2% of the parameters and >15× faster total inference. It reconstructs the diffusion backbone with a Local-λ Mix Interaction (LλMI) block (Local-λ + Interactive-λ modules) that summarizes spatial contexts and global semantic priors into fixed-size linear matrices, paired with an adaptive multi-granularity distillation strategy operating in latent space to avoid pixel-space decoding.
DragMesh-2 — dexterous hand-object interaction (HuggingFace, 63 upvotes). A contact-driven framework for multi-finger dexterous manipulation of articulated objects, where the target part can’t be directly actuated and motion must emerge through sustained hand-handle contact. Its PICA mechanism (physically informed contact-aware training) injects physical signals into policy learning without tactile/force feedback, improving robustness and task success under changing contact loads. Across seven GAPartNet objects it shows stronger robustness under contact-load variation than baselines.
Playful Agentic Robot Learning (HuggingFace, 37 upvotes). Introduces RATs (Robotics Agent Teams) for self-directed “play-time” skill acquisition before downstream tasks arrive: the embodied coding agent proposes novel-yet-learnable exploratory tasks, executes robot-code policies, verifies progress, diagnoses failures, retries with dense step-level feedback, and distills successes into a persistent code skill library. Play-learned skills improve held-out tasks by 20.6 points on LIBERO-PRO and 17.0 on MolmoSpaces over CaP-Agent0, and transfer to other Code-as-Policy agents (RoboSuite +8.9, real-world +8.8) without finetuning.
Multi-LCB (HuggingFace, 33 upvotes). Extends the contamination-aware LiveCodeBench (previously Python-only) to 12 programming languages by transforming Python tasks into equivalents while preserving LCB’s release-date filtering and evaluation protocol (auto-tracking future LCB updates). Evaluating 24 LLMs revealed Python overfitting, language-specific contamination, and substantial multilingual performance disparities.
Policy, safety & infrastructure specifications
Agentic Resource Discovery (Google + Microsoft). Google and Microsoft jointly announced Agentic Resource Discovery (ARD), an Apache 2.0 open specification for publishing, discovering, and verifying AI tools, skills, agents, and other resources across federated registries. GitHub’s Agent Finder launches as an implementation for Copilot.
Google’s AI Control Roadmap (DeepMind, The Batch via TLDR). Google published a framework for building and managing advanced internal AI that adds system-level security assurance even if alignment is perfect — incorporating sandboxing, endpoint security, and prompt-injection resistance. It treats internal agents as potentially misaligned, providing assurance even if alignment fails, with AI alignment as a primary (not sole) defense.
Field commentary & essays
Misc analysis pieces. “HTML is all you need” (12 Grams of Carbon) pushes coding agents to use HTML as the core creative medium for graphics/visual media. 404 Media covered Adrian de Wynter’s argument that we anthropomorphize interfaces too quickly — “if AI is sentient, then so is Age of Empires II” — since the same primitives can be built inside a strategy game. “That Untravell’d World” (Hyperdimensional) argues a new, higher-stakes, more political phase of AI governance has begun. OpenAI is also notably backing/funding Rust maintainers, underscoring how flashy AI demos depend on unflashy infrastructure (compilers, libraries, build tools). Boston Dynamics Spot robots were deployed for FIFA World Cup security/inspection. Google rolled out Gemini-powered Gmail summaries globally (free, paid, Workspace; Android/iOS/web), and opened pre-orders for a $99 Gemini-powered Google Home Speaker (its first new speaker in six years) for conversational/multistep commands.