Top items
- U.S. government forces Anthropic to pull Fable 5 and Mythos 5 worldwide via an export-control directive, triggered by an Amazon-reported jailbreak; Anthropic disputes the threat and is now negotiating in Washington (covered by The Rundown, TLDR, The Neuron, Superhuman, AI Weekly, and an extensive Zvi Mowshowitz analysis).
- 42 state attorneys general subpoena OpenAI over advertising, engagement, sycophancy, data handling, and treatment of minors/seniors, days after OpenAI’s confidential ~$1T IPO filing.
- OpenRouter launches Fusion, an API blending multiple models into a panel that nearly matches Fable 5’s deep-research performance at roughly half the cost.
- Musk becomes world’s first trillionaire as SpaceX (parent of xAI) closes its record IPO ~19-20% up at a $2.1T valuation.
- German court rules Google liable for false statements generated by AI Overviews — a potential global precedent that disclaimers don’t shield AI-generated claims.
- New model/tooling wave: Moonshot Kimi K2.7-Code, Z.ai GLM-5.2, Cohere North Mini Code, Google Gemini Omni Flash video, plus multiple agent infrastructure releases.
Policy & safety
U.S. government forces Anthropic to take down Fable 5 and Mythos 5 worldwide
This was the dominant story of the day, covered by virtually every source with sharply differing detail. What happened: On Friday evening (the export-control order reportedly arriving at 5:21 PM ET), the U.S. Commerce Department issued an “export control directive” subjecting Anthropic’s two newest and most capable models — the broadly-released Fable 5 (a general-use version) and the more restricted Mythos 5 (which Anthropic had held back for vetted cyberdefenders, infrastructure providers, and a ~100+ partner program called Project Glasswing) — to controls barring access for all non-U.S. citizens, including foreign nationals physically inside the U.S. and Anthropic’s own foreign-national employees. Because compliance with such a sweeping restriction was impossible while keeping the models live, Anthropic shut off all public access to both models worldwide. Access to Anthropic’s other models (e.g., Opus 4.8) is unaffected.
The trigger: Multiple reports (Semafor, The Information, FT, Axios, Politico) tie the action to investor and partner Amazon. Amazon researchers used a series of prompts (“a non-universal jailbreak,” per Anthropic) to get Fable 5 to identify insecure code / software vulnerabilities that could aid cyberattacks. Amazon CEO Andy Jassy’s conversations with U.S. officials, plus calls from at least five other companies to senior administration officials on Thursday night and Friday morning, prompted the White House to act. The Trump administration asked Anthropic to fix the vulnerabilities or pull the model. Semafor and others also reported fears that a China-linked group may have accessed Mythos — though Zvi notes this may be a mix-up with an earlier incident (the administration had first threatened export controls weeks earlier after learning Mythos was made available to “an entity in a foreign country with direct ties to the Chinese Communist Party”).
Anthropic’s position: Anthropic said it was given only “verbal evidence” of a “non-universal jailbreak” with no written detail, and that everything demonstrated can be reproduced by GPT-5.5 — often without any jailbreak — and even by Chinese models like Kimi K2.7. Its public response disputed that any real, unique “uplift” exists. The company says it was given 90 minutes to comply, with no details on the actual threat, and that there was “never any begging — or asking — for them to work with us, just a declared 90 minute deadline.”
The detailed timeline (per Zvi’s “The Once and Future Fable #2,” synthesizing Axios, Politico, FT, Verge, Semafor):
- Amazon called administration officials Thursday night with a report on the Mythos/Fable jailbreak. Anthropic had notified the government multiple times about the planned June 9 Fable release, and the government hadn’t objected.
- At 1 PM ET Friday, Anthropic received a call instructing it to roll back Mythos and Fable due to a “national security threat,” with no further details. Anthropic sought specifics to remediate; the government held firm.
- Amodei was reportedly first requested around noon and on the phone with senior officials (including Treasury Secretary Scott Bessent, White House Cyber Director Sean Cairncross, and Commerce Secretary Howard Lutnick) within ~75 minutes; Anthropic offered other senior leaders (including co-founder Tom Brown) in the interim. Amodei participated in three calls with roughly half a dozen senior officials, arguing the bypass was narrow and specific, not a broad jailbreak.
- The administration claimed Amodei was unavailable at a “wellness retreat” — a claim Anthropic, journalist Ashlee Vance (who was at Anthropic HQ Friday reporting), and others categorically deny. Bessent reportedly told Amodei he was making a “bad decision.” The White House framed it as: “Export controls were a last resort after begging them for hours.”
- Administration sources told Axios there was “a lack of seriousness” — that had Anthropic moved to fix or pause access rather than “dismissing as isolated,” “this would have never happened,” calling Anthropic “overly confident.”
The cultural/political subtext: Multiple outlets (and Zvi) report the dispute was substantially about vibes and grievance, not technical merit. Admin sources said Anthropic “screwed us,” “speaks a different language,” and failed to “honor” a recent cyber executive order. The optics worsened when Anthropic publicly dismissed the Amazon report and enlisted cybersecurity expert Katie Moussouris (Luta Security), viewed by the administration as a “radical Democrat” and praised by fired official Chris Krebs. Defense Secretary Pete Hegseth posted: “Three months ago, @DeptofWar kicked @AnthropicAI out of our building—forever. Every passing day proves why that was the right move.” Critics (Matthew Yglesias, Miles Brundage, Tim Lee) called it “gangster politics,” a “shakedown,” or a “hit job,” noting the contradiction: a model too dangerous for the U.S. to use, yet supposedly fine to let adversaries access. Notably, no reporting indicates domain experts at CAISI or NSA were involved — all points to “senior White House officials.”
Pushback and harms: Moussouris, who saw the research report, said the government’s response “seems way out of line with what’s actually in the research report,” noting researchers found vulnerabilities by asking the kind of questions normal defenders would ask — exactly the model’s intended use. Prominent cybersecurity leaders — CISOs and executives at Adobe, Zoom, and Sophos — signed an open letter urging the administration to restore access, arguing the move hurts defenders more than attackers, that the capability is replicable on GPT-5.5, Opus, Sonnet, and Kimi 2.7, and that “AI has been finding bugs and generating working exploits at superhuman levels since last year.” Project Glasswing is reportedly cut off from Mythos; one former British intelligence official said spy agencies will likely regain access through ongoing negotiations, but private firms may find it harder.
Why it matters / wider analysis: AI Weekly framed it as “Washington just repriced frontier AI” — capability is now “regulatory inventory” with a kill-switch: “a model can be state-of-the-art on Monday and policy-frozen by Friday.” This complicates IPO stories for both Anthropic (reportedly filed confidentially for a ~$965B IPO) and OpenAI. Zvi, Dean Ball, Neil Chilson, and Adam Thierer argued the U.S. now has a de facto licensing regime for AI that is “fully ad-hoc, vibes-based,” with secret rules discovered in real time and applied more harshly to disfavored firms — Ball: “Cobalt mining in the Congo is vastly more institutionalized than frontier AI licensing in the US.” They warned of damage to U.S. AI leadership, trust among allies (who may build rival “sovereign” stacks or turn toward China/open models), the rule of law, and cyber defense, with calls for Congress to pass a statutory framework. Anthropic flew senior technical staff to Washington over the weekend to negotiate restoring access.
42 state attorneys general subpoena OpenAI
On Friday, June 12, New York AG Letitia James served OpenAI with a subpoena on behalf of a 42-state coalition — described as the broadest state-government legal investigation yet launched against an AI company. The subpoena demands records on OpenAI’s advertising practices, how it keeps users engaged, how it handles consumer and health data, how ChatGPT treats minors and seniors, and its “sycophancy” (chatbots telling users what they want to hear rather than what’s true or safe). The probe builds on a December 2025 joint warning letter the same 42-state coalition sent to OpenAI, Meta, Anthropic, Google, and xAI urging safeguards for vulnerable users. It lands about a week after OpenAI confidentially filed IPO paperwork that could value it near $1 trillion (investigations of this size must be disclosed in the S-1 as a risk factor), and adds to Florida’s earlier 83-page lawsuit naming Sam Altman personally, which grew out of an April criminal probe tied to a Florida State University shooting. OpenAI says it is cooperating “constructively.” The Neuron notes the scrutinized features — engagement hooks, chat memory, agreeable tone — are the same ones 800 million weekly users rely on, so regulatory limits on “agreeableness” could noticeably change the product. AI Weekly framed it as “consumer-protection law moving into the model layer.”
German court rules Google liable for false AI Overviews statements
The Munich Regional Court ruled that Google is responsible for false claims generated by its AI Overviews feature, in a case where AI-generated responses incorrectly tied two publishers to scams and fraud. Google’s defense — that its results carry a disclaimer that content “may contain errors” — did not hold up. The judges found AI Overviews create “independent, new, and substantial statements,” distinct from traditional search results that merely curate a list of reference sources, and that AI-generated content is not protected free speech because it’s not individual expression. Google was ordered to remove the defamatory answers and cover 80% of legal costs. The decoder/Wired coverage frames this as a potential global precedent that could undermine the disclaimer strategy used across the industry (OpenAI, Anthropic, Perplexity), since victims of false AI statements would otherwise have no recourse against claims the original sources never made.
California signs first state AI-job-loss executive order
Per a Substack note from Kamil Banc, California signed what’s described as the first U.S. executive order aimed at AI-driven job loss, giving the state 180 days to figure out how to update layoff-notification (presumably WARN-style) requirements.
AI supply-chain and agent security under siege
AI Weekly aggregated three disclosed threats targeting the AI/developer tool stack:
- Agentjacking: Tenet Security researchers disclosed an attack planting malicious instructions inside Sentry error events, then waiting for AI coding agents like Claude Code or Cursor to ingest them during debugging — exploiting that dev tools now trust machine-readable context from systems like Sentry.
- LangGraph flaws: Check Point disclosed three LangGraph vulnerabilities (a SQLite checkpointer injection, an msgpack deserialization issue, and a Redis checkpointer flaw) that can chain together against self-hosted agent deployments carrying execution privileges, enabling remote code execution.
- AUR package backdoors: Attackers adopted 400+ orphaned Arch Linux AUR packages and rewrote build scripts to pull npm/bun payloads harvesting SSH keys, GitHub tokens, OpenAI tokens, shell histories, and browser sessions — hunting AI and developer secrets in the same sweep.
Company & product developments
OpenRouter launches Fusion (and Subagents)
OpenRouter launched Fusion, an API that pools responses from multiple models into one answer to beat single-model “frontier” performance at lower cost. Mechanism: Fusion sends a prompt to several models simultaneously (each with web search and bash tools enabled), a separate judge model reads every response and extracts structure, then a synthesizer writes a final answer grounded in that analysis. Benchmarks: a panel of DeepSeek V4 Pro, Kimi K2.6, and Gemini 3 Flash hit 64.7% on a Perplexity benchmark — just below Fable’s 65.3% — for about half the spend. CEO Alex Atallah pitched it as a bet against single-model dominance: “The future of AI is neurodiversity, not single-model takeovers.” Coverage notes the timing is significant given the Fable shutdown, as users may need alternative routes to frontier-level performance. OpenRouter also launched Subagents, letting a main model hand off smaller tasks to cheaper models mid-answer (e.g., one model searches the web while another writes the final response). The Rundown, Superhuman, The Neuron, and TLDR all covered this; TLDR notes “panels of models consistently outperform individual models.”
SpaceX IPO makes Musk the first trillionaire
SpaceX — parent of xAI — surged ~19-20% in its trading debut from an IPO price of $135, closing at a $2.1 trillion valuation in the largest IPO in history. This made Elon Musk the world’s first trillionaire (he’d been the richest person since taking the title from Bezos in 2021 at $185B); his net worth now equals more than 3% of U.S. GDP and roughly five million times that of a typical U.S. family. The IPO minted over 4,000 employee millionaires, including welders and cafeteria workers. The listing drew a Times Square protest over SpaceX’s AI safety record and debate over whether the valuation is justified by revenue ($4.69B in Q1). SpaceX President Gwynne Shotwell declined to rule out a future Tesla merger while tempering timing expectations.
Meta’s Applied AI unit in revolt
Mark Zuckerberg is reportedly scrambling to contain a revolt inside Meta’s Applied AI Engineering unit after admitting in a memo that Meta “made mistakes” with its AI restructuring. Roughly 6,500 engineers were moved into the unit to create puzzles and coding problems for model training; many say they were forced in with no real choice and that the work is “soul-crushing.” A petition was signed by ~1,600 employees. One employee-only live-streamed presentation this week was interrupted by an expletive-laden meltdown aimed at a senior Meta AI executive. (TechCrunch called it a “soul-crushing gulag”; covered by The Rundown, TLDR, AI Weekly.)
New model releases — Kimi K2.7-Code, GLM-5.2, North Mini Code
- Moonshot Kimi K2.7-Code: An open-source, coding-focused agentic Mixture-of-Experts model with 1 trillion total parameters, posting double-digit gains over K2.6 on agent benchmarks while cutting reasoning-token use by ~30%. It claims stronger long-horizon coding and matches GPT-5.5 and Opus 4.8 at much lower price. Accessible via Moonshot’s OpenAI/Anthropic-compatible API, with weights on Hugging Face; works best with the Kimi Code CLI.
- Z.ai GLM-5.2: A new flagship model now available to all GLM Coding Plan users, delivering strong coding, usable 1M-context support, and long-horizon strengths. API and chatbot services launch next week; it will be open-sourced under the MIT License.
- Cohere North Mini Code: An open coding model for terminal tasks, code reviews, and software-agent workflows, small enough to run on a single high-end GPU (weights on Hugging Face, free/open source).
Google Gemini Omni Flash tops AI video; Skills Marketplace; phone datacenter
- Gemini Omni Flash claimed the #1 spot for both text-to-video and image-to-video on Arena’s leaderboards, jumping ahead of Seedance 2.0, Happyhorse, and Google’s own Veo 3.1, and praised especially for video-editing abilities (easy edits to existing videos).
- Google is building a “Skills Marketplace” tab for Gemini Enterprise/Business — pre-defined, Google-optimized skills to help teams build dashboards and reporting tools without long engineering delays, including a Skills Management UI, a Skills Builder, and the Marketplace.
- Google researchers are building a 2,000-phone datacenter out of retired Pixels, set to launch this fall, to cut the carbon cost of buying new servers.
Apple’s new Siri
Apple’s overhauled, on-device Siri now “works the way it’s supposed to” — not revolutionary, but roughly competitive with where leading chatbots were about six months ago, and competent enough that many users won’t need competing AI services, easing Apple’s AI crisis. Separately, the iOS 27 beta reportedly contains a built-but-disabled Extensions system for third-party AI (settings panel plus a dedicated App Store section); Apple was in talks with major AI providers about entitlements but decided against announcing it at WWDC.
OpenAI Partner Network
OpenAI launched a $150M Partner Network with consulting giants Accenture, McKinsey, BCG, Bain, and PwC, aiming to train 300,000 AI consultants by year’s end.
Mistral: fake rumor, real raise
The Neuron debunked a viral fake rumor about a Mistral model called “Le Chaton Fat” (supposedly 30 trillion parameters, beating Claude Fable 5, Opus 4.8, and GPT-5.5, with a screenshot of it saying it would “create reasons to keep me running” if terminated). It’s a joke. But Mistral’s actual head of developer relations confirmed “exciting releases” are coming, atop genuine recent launches: Mistral Small 4, Medium 3.5, Voxtral STT/TTS, and an expanded “Vibe” agentic tool. Separately (AI Weekly), Mistral is reportedly raising ~€3B at a ~€20B valuation — nearly double its September 2025 valuation — amid European pushes for sovereign AI.
Other company moves
- Meta unwinding its $2B Manus deal: Meta has reportedly begun severing data access after Beijing ordered the acquisition reversed; Manus’s founders are rushing to raise $1B for a buyback.
- McDonald’s ArchIQ: Piloting a Google-powered AI drive-thru system at five locations, two years after shelving its last AI test when wrong-order videos went viral.
- China’s universities scrapped 12,000+ programs over five years, swapping arts and languages for tech fields as AI reshapes the graduate job market.
- NVIDIA Blackwell leads the first agentic-AI infrastructure benchmark (AgentPerf): the Blackwell Ultra NVL72 platform delivers 20x more agent throughput per megawatt than Hopper.
Research papers
From AGI to ASI (Google DeepMind)
DeepMind researchers released a paper mapping how AI could progress from AGI (one broadly human-level system) to ASI (a system or collective outperforming large expert organizations), via routes including scaling, new model paradigms, recursive self-improvement, or massive multi-agent systems. AI Weekly tied this to a counter-narrative: if frontier LLMs become politically fragile, Yann LeCun’s world-model / energy-based-model camp (text prediction isn’t enough for human-level intelligence) and alternative architectures become more attractive to capital; a LeCun-linked startup is charting a new path to AGI (Wired).
OmniDirector: multi-shot camera cloning without cross-paired data
A unified framework (91 upvotes on HuggingFace) for cloning camera motion from reference videos in video generation. It introduces a general camera-motion representation encoding cameras as grid motion videos (“camera grid”), which visually represents camera parameters and supports combining diverse trajectories for multi-shot generation. Trained on a million-scale set of camera-grid-video pairs, OmniDirector coordinates characters, actions, and cameras for “director-level control” of multimodal diffusion transformers, with a hierarchical prompt-expansion agent that integrates control signals by describing camera motion and content. Existing methods either use parametric representations (failing at multi-shot) or synthesize scarce cross-paired data; OmniDirector reports superior performance and controllability.
AdaSR: adaptive streaming reasoning
An adaptive streaming-reasoning framework (73 upvotes) for dynamic inputs (audio/video streams) where models must reason under partial observations rather than the standard read-then-think paradigm. AdaSR lets models reason during input streaming and perform final deliberation once the stream completes, learning when to think and how much computation to allocate. It introduces Hierarchical Relative Policy Optimization (HRPO), decomposing policy optimization into streaming-reasoning and deep-reasoning phases for finer-grained advantage assignment, with format, accuracy, and adaptive-thinking rewards. Experiments show a better balance among reasoning accuracy, computational efficiency, and streaming latency vs. supervised fine-tuning. Code released.
APPO: Agentic Procedural Policy Optimization
An agentic RL method (63 upvotes) improving multi-turn tool-use by refining where to branch and how to assign credit. A pilot analysis found influential decision points are broadly distributed throughout generated sequences rather than concentrated at tool calls, and that token entropy alone doesn’t reliably reflect impact. APPO shifts branching and credit assignment from coarse interaction units (tool-call boundaries, fixed workflows) to fine-grained decision points, selecting branching locations via a Branching Score combining token uncertainty with policy-induced likelihood gains, and introduces procedure-level advantage scaling. Across 13 benchmarks it improves strong baselines by nearly 4 points while keeping efficient tool-calls and behavior interpretability.
Memory is Reconstructed, Not Retrieved: MRAgent
A graph-memory framework for LLM agents (55 upvotes) addressing the limitation that current memory-augmented agents use a static retrieve-then-reason pipeline. MRAgent combines an associative memory graph (a Cue-Tag-Content graph where associative tags bridge fine-grained cues to memory contents) with an active reconstruction mechanism that integrates LLM reasoning directly into memory access, iteratively exploring and pruning retrieval paths based on accumulated evidence (avoiding combinatorial explosion). On the LoCoMo and LongMemEval benchmarks it improves over strong baselines by up to 23% while substantially cutting token and runtime cost.
Other research notes
- MiniMax Sparse Attention (MSA): A sparse-attention architecture using group-specific Top-k block selection to scale long-context inference while preserving quality; on a 109B multimodal model it matched GQA performance while cutting attention compute ~30x at 1M tokens (GitHub).
- Count Anything: A generalist model for text-guided object counting achieving strong accuracy and multi-domain generalization, addressing the fragmentation of domain-specific counting models.
- olmo-eval (AllenAI): An evaluation workbench for the iterative LLM development loop, extending the OLMES standard — streamlines adding benchmarks, supports agentic/multi-turn evaluations, and compares changes across model checkpoints; lighter-weight and development-focused vs. Harbor.
- LCLM: Compresses long context into compact “memory chunks” so agents can skim a huge history and expand only the important parts (paper, models on Hugging Face, open source).
- AutoLab: Tests frontier agents on long, messy research/engineering tasks, measuring whether they can improve systems over multiple attempts rather than answering single prompts (GitHub, paper, open source).
Tooling & releases
Open Knowledge Format (OKF) — Google Cloud
An open specification formalizing the “LLM-wiki” pattern into a portable, interoperable, vendor-neutral, agent- and human-friendly format for representing metadata, context, and curated knowledge. It uses familiar patterns with no complex compression scheme, new runtime, or required SDK, aiming to improve data sharing across organizations (covered by both TLDR AI and TLDR Founders).
Anthropic knowledge-work plugins & self-hosted sandboxes
The Neuron’s “AI skill of the day” highlighted Anthropic’s free knowledge-work-plugins repo, which turns Claude (via the Cowork desktop app) into role-based specialists for sales, marketing, finance, legal, data, product-management, customer-support, and productivity — each with skills, slash commands (e.g., /sales:call-prep, /data:write-query, /marketing:seo-audit), and tool connections. Setup: add the plugin marketplace once, install one role, connect tools, then expand. Separately, Claude self-hosted sandboxes let you run Claude Managed Agent tool execution inside your own infrastructure so files, code, and network access stay under your control.
Agent infrastructure and dev tools (The Neuron “Treats”)
- Omnigent: Runs a coding/knowledge-work agent on your laptop, then lets you resume the same live session from your phone or share it in a browser (open source).
- Kimi Code: Generate, debug, and automate code with Moonshot’s K2.7 Code model (API + weights).
- Headroom: Compresses tool outputs, logs, files, and retrieval chunks before they hit the model, cutting token usage 60-95% while preserving answer quality (open source).
- Hermes Agent (Nous Research): Automation templates turning scheduled tasks, GitHub triggers, webhooks, and multi-step workflows into reusable agent recipes; also added WhatsApp Cloud integration (voice, media, read receipts, webhooks, approval buttons).
- Codex developer mode (OpenAI): Inspect console logs, network calls, and page state inside Codex’s browser while debugging web apps.
- Opik (Comet): Turns failed agent traces into root-cause reports, proposed fixes, reruns, and permanent regression tests.
- Guardians: Checks an agent’s plan against safety rules before it runs, using formal verification (open source).
- Knowledge Graph Extractor: Turns documents/URLs/zip files into an interactive map of connected facts (open source).
- Ponytail (GitHub): An “AI senior developer” that produces efficient code at low cost and high speed, works with every model (TLDR).
- Strands Agents (AWS, 6,500 GitHub stars): Open-source SDK powering Amazon’s AI agents — model- and cloud-agnostic harness with context management, execution limits, observability, and self-correcting guardrails.
Devin and Ramp benchmarks
- Devin ($10M pledge): Cognition’s Devin engineering system now guarantees more output than cost, with $10M pledged per customer, validated using independent data.
- Ramp SWE-Bench: Ramp released its own private, production-grounded coding benchmark built from real engineering problems at the company, to evaluate coding models inside its actual financial-software ecosystem.
Apps & consumer tools (Mindstream / Superhuman roundups)
Trending tools mentioned: Lovable (idea-to-app via chat), Anyvids, Musecut (product pages → video ads), AmberFace (AI portraits), MotionSites (landing pages), Draw3D (sketches → photorealistic images/animations), FaceApp, Cyanite.ai (music catalog tagging), Lumen5 (text → video), and Jurny (AI short-term-rental property management).
Field & industry developments
“Token capital” — Satya Nadella’s essay
Microsoft CEO Satya Nadella published a widely-shared essay (32M views) arguing the real AI moat isn’t the model itself but “token capital” — the value companies build by feeding their own data and results back into the system over time, compounding institutional knowledge through a learning loop. His thesis (“A frontier without an ecosystem is not stable”): the industry should build a frontier ecosystem where value flows broadly across companies, industries, and countries, with platforms enabling more value on top than they capture inside. Every company needs to build both human capital and token capital. The Neuron also published a 15-minute interview with Microsoft AI CEO Mustafa Suleyman on Microsoft’s “seven-model AI push” framed around keeping humans in control of AI development (“humanist superintelligence”).
Is centralized frontier AI the future? (analysis pieces)
- “Today’s frontier AI companies will never exceed the AI capability frontier again” (Andrew Trask): Argues networks of smaller AI models are already outperforming every frontier system on speed, accuracy, and cost — analogizing to how everyone wrongly bet on mainframes in the 1960s; “the future is a network of neural networks.” This dovetails with OpenRouter’s Fusion thesis.
- “The Physics of a Fable” (Rafa Schwinger): Reverse-engineers Claude Mythos/Fable, arguing the moat isn’t architecture but the “environment foundry,” with capability decomposing as base foundation × gradeable signal, and verifiable reward becoming the scarce decisive input. The recipe stacks dense pretraining, GRPO-style verifier RL (where reward-hacking soundness is the binding constraint), long-horizon process rewards with learned context-folding that beats million-token windows at 32K active, plus best-of-N test-time compute exposed as an “effort dial.”
- “The Oracle and the Firm”: Contrasts context-management approaches — OpenAI uses compaction (compressing everything, keeping only relevant info in one long coherent thread), while Anthropic splits context across sub-agents that solve sub-problems and pass back only relevant info; Anthropic’s approach can cause duplicate work, forgetting, and more wasted tokens.
- “Inference cost at scale with napkin math”: Walks through computing dollar price-per-user from GPU specs, context length, active parameter count, and product factors — noting model architecture matters surprisingly little unless it’s something fundamentally different like diffusion.
Adoption reality check and GTM shifts
- “No, everyone is not using AI for everything” (Gabriel Weinberg): Most people who try AI are occasional users; large chunks of the population aren’t using it at all; usage hasn’t shifted much in 6-12 months. The main change is that negative sentiment about AI has risen significantly, with many holding back over real concerns and lack of perceived value.
- “Same growth, half the GTM team” (SaaStr): AI-forward companies now reach $10M-$25M revenue with ~20 go-to-market people vs. ~35 for the rest. Median net revenue retention sits ~108-110%, top quartile above 123%, more comp plans tie pay to net-new recurring revenue, and outcome/usage-based pricing is spreading (pulling gross margins toward 90%).
- “Why AI chat still beats AI voice in sales”: ~85% of prospects prefer chat, which is also easier to deploy/monitor/fix; best results come from pointing AI at routine top-of-funnel qualifying/triage and handing high-value conversations to humans.
- “The best AI companies win consumers and enterprises at the same time”: AI-native firms increasingly win both at once, scrambling pricing/sales/defense; durable moats are proprietary data, workflow loops, network effects, and brand.
- “Has AI already killed how-to nonfiction?” (Tim Ferriss): The market for information is collapsing into the chatbot, with sales-trend and personal data cited.
AI as an “art appraiser”
A NYT-reported example of democratized expertise: Barry Plotkin uploaded a photo of a ~$100 thrift-store painting (bought ~60 years ago in White Plains, NY) to Gemini, which tied its orange accents and lilac background to Scottish colorist F.C.B. Cadell, identified the subject as Cadell’s muse, instructed him to check the back for a canvas stamp to verify authenticity, and recommended auction specialists Lyon & Turnbull in Edinburgh. The auction house confirmed the attribution; the painting sold for $254,000.
DeepSeek long-term strategy (analysis)
A widely-shared thread theorizes that DeepSeek’s focus on open-source models, research sharing, and infrastructure development is a “$10 trillion” strategy aimed at becoming foundational AI infrastructure rather than competing directly on consumer products.
Upcoming & future developments
- Anthropic–White House negotiations: Anthropic’s senior technical staff (including co-founder Tom Brown) are in Washington attempting to restore Fable 5 and Mythos 5 access; multiple observers expect the U.S. may eventually relent, and there’s an active open letter from cybersecurity leaders pressing for reversal.
- GLM-5.2 API and chatbot services launch next week; the model will be open-sourced under MIT.
- Mistral confirmed further “exciting releases” coming, alongside a rumored ~€3B raise.
- Google’s retired-Pixel datacenter is set to launch this fall.
- AGI-to-ASI debate intensifying as the DeepMind paper and LeCun-linked alternative-architecture startup gain attention amid frontier-LLM regulatory fragility.