AI Daily Digest

Tuesday, June 30, 2026

4,536 words · All issues

Top items

  • Meta’s Brain2Qwerty v2 decodes whole sentences from non-invasive brain scans at 61% average word accuracy (78% best), a leap from the ~8% of prior non-invasive methods, with code and v1 dataset open-sourced.
  • California signs a first-of-its-kind deal with Anthropic to put Claude in all state agencies, cities, and counties at a 50% discount — even as the Pentagon recently flagged Anthropic as a “supply-chain risk.”
  • Four major workplace studies, read together, show AI productivity gains are real but lopsided: novices +34%, experts −19%, and 95% of enterprise GenAI pilots returning no measurable P&L.
  • Chinese open-weight models surge: Meituan open-sources 1.6T-parameter LongCat-2.0 (first trained end-to-end on domestic chips), Zhipu’s GLM-5.2 matches top US models on cybersecurity bug-finding, and Grok 4.5/V9 enters private beta.
  • Cursor and OpenClaw go mobile; coding shifts from typing to supervising agents from your phone.
  • Anthropic Claude goes GA on Microsoft Foundry/Azure on NVIDIA GB300 Blackwell Ultra — Anthropic’s first-ever NVIDIA deployment — while Amazon weighs cheaper alternatives after Anthropic’s price hikes.

Research papers & benchmarks

Meta Brain2Qwerty v2 — non-invasive brain-to-text reaches full sentences. Meta introduced version 2 of its Brain2Qwerty system, which decodes typed sentences in real time from non-invasive magnetoencephalography (MEG) brain recordings, a major step beyond v1, which could only spell things out one character at a time. Nine volunteers each sat about 10 hours inside a brain scanner typing while the system read along, producing roughly 22,000 sentences of training data (Meta says v2 is trained on ~10× more data than v1). The architecture uses two models: one reads the raw brain signals as people type, while a second adds semantic meaning. v2 reached an average word accuracy of 61% — versus the ~8% highs of the best prior non-invasive rivals — with the top-scoring volunteer hitting 78%, accuracy that begins to approach setups that previously required surgical implants. Meta reported accuracy climbs with more data, suggesting the remaining gap with surgical BCIs “could be further narrowed through data scaling alone.” Meta open-sourced the training code for both v1 and v2 plus the v1 dataset on Hugging Face. The stated goal is to help people who can no longer speak after a stroke, accident, or brain disorder; because surgical implantation has been a hard barrier to mass adoption, a non-invasive path could broaden who benefits, and open-sourcing means other labs can push the same curve. (Reported by The Rundown, TLDR, The Neuron, Superhuman.)

Anthropic Economic Index — round-the-clock Claude usage. Anthropic published the latest report in its Economic Index, pairing a new continuous (hourly) usage-sampling setup — replacing the seven-day slices used in past reports — with a survey of 9,700 users. The hourly data shows news questions peaking in the morning, recipes at dinner, sleep advice before dawn, gardening in the afternoon, sermon-related queries before sunrise, and chat spikes on dates like Tax Day. Personal (non-work) Claude chats rose from roughly one-third of usage during the workweek to nearly half on weekends, skewing toward health, money, and emotional-support prompts. Notably, users who delegated more work to Claude also expected AI to handle more of their tasks next year and reported feeling better about their income, career stability, and sense of purpose. Anthropic frames the granular data as evidence that AI is assimilating into everyday schedules, not just office workflows, making habitual trust and use as telling as benchmark scores.

New benchmarks and inference research. Several evaluation and efficiency papers landed: DeepSeek open-sourced DSpark, a framework that speeds up LLM inference by up to 85% without changing model outputs — it acts as a “scout” running a few steps ahead, guessing the likely decoding path and letting the larger model quickly verify which steps are safe, accelerating when guesses are good and avoiding wasted checks when they’re weak (VentureBeat). Cognition’s Devin Fusion (see Tooling) details its routing math. RoadmapBench evaluates long-horizon agentic software development across real version upgrades in 17 repositories, with 115 tasks requiring a median modification of 3,700 lines across 51 files (arXiv). Artificial Analysis launched AA-Briefcase, a private benchmark for long-horizon agentic knowledge work across spreadsheets, presentations, and memos. DiScoFormer is a single transformer that estimates both density and score from data in one pass without retraining, using cross-attention to adapt to new distributions on the spot; trained with Gaussian Mixture Models, it cut score error 6.5× and density error more than 37× versus classical kernel density estimation in 100 dimensions (Hugging Face). Nvidia’s CHORD project demonstrated transferring human dexterous-manipulation skills to robot policies using contact-focused demonstrations. The essay “RL Beyond the Verifiable” argues the next leap in reinforcement learning will come from bringing RL’s success in verifiable domains to harder-to-verify tasks, surveying why verifiability is the constraint and which companies are attacking it (tanayj.com).

Tooling & releases

Cursor for iOS — coding agents from your phone. Cursor brought its agentic coding platform to iPhone and iPad with a new app (public beta), letting users kick off agents or take over agents already running on their PC from anywhere — shifting development from desk work toward on-the-go supervision. Users pick a model and start an agent by voice or slash command, running it on Cursor’s cloud or their own machine and moving between the two; Live Activities and push alerts hit the lock screen the moment an agent finishes, stalls, or has a pull request ready to review and merge from the phone. Use cases include incident handling, resolving customer issues, and acting on user feedback. The launch trails mobile pushes from Anthropic and OpenAI; Claude Code lead Boris Cherny said “Most of my coding now is on my phone.” A viral release video drew 3M views. Context: after an April partnership, SpaceX this month exercised a reported $60B deal to acquire Cursor, on the heels of SpaceX’s record IPO, and Cursor is preparing its first post-merger Opus-sized model. (The Rundown, TLDR, The Neuron, Superhuman.)

OpenClaw goes mobile. OpenClaw — the platform that popularized personal AI assistants in January — launched native iOS and Android apps, letting users command their AI assistant (for emails, calendars, home automation, and more) from a phone; free to try. (Superhuman, The Neuron.)

Cognition’s Devin Fusion. Cognition launched Devin Fusion in preview, a multi-model “harness” using a dual-agent architecture: a main frontier-model agent paired with a cheaper “sidekick” agent for dynamic model routing, optimizing task handling and avoiding costly cache misses. It cut costs 35% on the FrontierCode benchmark while maintaining top-tier performance, rising to a 41% cost reduction with Fable 5 integration. (TLDR AI, The Neuron, The Rundown.)

Other launches and tools. OpenAI’s Codex hardware (“Codex Micro”) was displayed at the AI Engineer World’s Fair — a keyboard designed to supercharge Codex usage, built with accessories company Work Louder (The Verge). OpenAI is also rolling out a bidirectional voice mode for ChatGPT that can speak, hear, and listen simultaneously (and beatbox) (Testing Catalog). Gemini’s personalized AI image generation (“Nano Banana”-powered) is now free for all eligible US users, generating images from the model’s understanding of users’ preferences without explicit prompting; the opt-in “Personal Intelligence” feature lets users choose which apps Gemini can access, and Google has further updates planned including a “Daily Brief,” a revamped interface, the Gemini Omni video model, and a personal agent called Gemini Spark (TechCrunch). Figma beefed up its AI design agent with custom tools, context, and skills, connecting to Hex, Notion, and Atlassian from the Figma dashboard. ByteDance’s Seedance 2.5 video model can create 30-second clips from one prompt and accepts up to 50 reference attachments for greater control (CNET). Sakana Fugu launched after a Claude ban, with Fugu Ultra scoring 93.2 on LiveCodeBench (beating Fable) starting at $5 per million input tokens. Base1 is Base44’s new in-house app-building model. ClinePass gives discounted, key-free access to GLM, Kimi, DeepSeek, MiMo and other open coding models inside Cline for $9.99/month. Mistral Workflows is a durable, fault-tolerant orchestration platform for multi-agent pipelines. OpenAI’s Record & Replay tool (macOS) lets you teach Codex any repetitive desktop task by recording your screen, then turning it into a reusable, schedulable skill. Several context/assistant tools also shipped: Draft (captures meeting/Slack/GitHub context to inject into agent sessions), Halo (private iPhone assistant unifying mail, calendar, reminders, music, health), and Bloome (multiple AI agents in group chats alongside humans). National Design Studio released Rampart, an on-device tool to keep personal info separate from AI chatbots.

Company & product developments

California’s landmark Anthropic deal. Governor Gavin Newsom announced a first-of-its-kind partnership giving California state agencies — plus cities and counties — access to Claude at half price, paired with free Anthropic workforce training and technical assistance. Claude will be the first AI productivity tool offered through the state’s new SITeS portal and the first AI cleared by the state; the DMV, the Department of Healthcare Services (the country’s largest Medicaid agency), and the Office of Emergency Services are already using it. The deal stands in sharp contrast to the Pentagon’s recent designation of Anthropic as a “supply-chain risk,” signaling a divergent US state-versus-federal path on frontier-model adoption. (TechCrunch via AI Weekly Alerts, The Rundown.)

Anthropic Claude GA on Microsoft Foundry / NVIDIA. Anthropic, Microsoft, and NVIDIA announced that Claude Haiku, Sonnet, and Opus are now generally available in Microsoft Foundry on Azure, running on NVIDIA GB300 Blackwell Ultra NVL72 systems with Quantum-X800 InfiniBand — Anthropic’s first-ever deployment on NVIDIA hardware. It builds on their November tripartite strategic partnership and includes NVIDIA Verified Agent Skills and a Secure Agent Workspace reference design for enterprise agentic AI. Early Microsoft benchmarks cite a 40% token-generation speedup for Claude Sonnet on GB300 versus H100. (NVIDIA blog, The Neuron.)

Amazon weighs cheaper alternatives to Anthropic. After Anthropic renegotiated its AWS contract to shift to per-token pricing — which could substantially increase Amazon’s costs — Amazon is reportedly weighing a move to OpenAI models and its own Nova models to cut spending (The Information, TheNextWeb). This dovetails with a broader cost-control wave: companies like Uber and ServiceNow blew through their entire 2026 AI budgets within months, prompting many to consider shifting workloads to cheaper open-weight models that increasingly match proprietary performance, though performance, security, and engineering-overhead trade-offs mean open-weight isn’t right for every team (Superhuman, 404 Media’s “Tokenpocalypse” via AI Weekly).

Claude expands across the enterprise stack — with friction. Claude became generally available in Microsoft Foundry (above) and Anthropic launched Claude Tag, which monitors activity in a work messaging app and, with preset guidance, sends alerts about posts that may impact the user’s day, drops comments in conversations, and can fix code issues (Bloomberg framed it as wanting Claude to be “your new Slack coworker”). The product drew skepticism over unclear benefit versus existing Slack/Teams alerts. Notably, Salesforce helped promote Claude Tag inside Slack, confusing employees because Slack already has Slackbot and the Agentforce platform (which itself runs on Claude); Salesforce expects to spend $300M on Anthropic tokens this year and holds roughly a 1% stake in Anthropic, giving it financial reasons to promote Claude Tag despite the competitive tension (TheNextWeb).

Getty–OpenAI image deal. Getty Images signed a deal in which OpenAI pays Getty to display Getty images in ChatGPT search results — a notable lifeline given Getty’s stock has fallen from a high of ~$30 to under $1, and an encouraging sign that culturally significant, irreplaceable IP may retain value in the AI era (Yahoo Finance).

Apple talent and acquisitions. Apple’s Vision Pro and smart-glasses chief Paul Meade is leaving for OpenAI’s hardware unit, reuniting with ex-Apple colleagues Jony Ive and Tang Tan to build AI-powered devices (Bloomberg). Separately, Apple acquired Rabbit 3 Times, maker of “Play,” a visual Swift development tool that won an Apple Design Award; the acquisition appears aimed at asset-stripping for incorporation into Apple Creator Studio, with the app gone and the company’s website taken down (AppleInsider).

Other company moves. Arena reached a $100M annual revenue run rate eight months after launching its real-world AI evaluation product. Google Cloud will sell specialist scientific AI models from SandboxAQ — large quantitative models trained on scientific equations and laboratory data — widening access for drug discovery, materials science, and semiconductor manufacturing; researchers can combine them with Gemini (Gemini for reasoning/interface, the quantitative model for the underlying science). Rocket Lab is acquiring satellite operator Iridium for $8 billion to challenge SpaceX, gaining control of Iridium’s 66-satellite low-Earth-orbit fleet, valuable spectrum rights for direct-to-device connectivity, and a network serving 2.5M+ users; Rocket Lab plans to develop an upgraded constellation to replace the current one (TLDR, The Verge). Mercury launched Mercury Command, a plain-language AI financial interface across its platform (payments, forecasting, categorization, invoices) with traceable reasoning and human approval gates.

Field & industry developments

AI productivity is real but radically uneven — special analysis. AI Weekly synthesized four of the most-cited workplace studies of the AI era into one conclusion: gains bend toward the inexperienced, the well-scoped, and the verifiable — not toward experts doing hard, familiar work, and not toward enterprises buying pilots.

  • Brynjolfsson, Li & Raymond (NBER/QJE, “Generative AI at Work”): Across 5,179 customer-support agents, average productivity rose 14%, but the gain was almost entirely carried by the least-experienced workers (+34%) while veterans barely moved. AI worked by encoding the best agents’ tacit know-how and handing it to newcomers, compressing months of learning into a prompt.
  • Dell’Acqua et al. (HBS/BCG, “Navigating the Jagged Technological Frontier”): BCG consultants using GPT-4 completed 12.2% more tasks, 25% faster, with higher quality — but only inside AI’s “jagged frontier.” On a task deliberately chosen to sit just outside it, AI-assisted consultants were 19% less likely to be right, because the tool’s fluent confidence is identical on both sides of the invisible line. Below-average performers gained most, narrowing the gap between weak and strong consultants.
  • METR randomized trial (arXiv 2507.09089): 16 experienced open-source developers on 246 real tasks in repositories they maintain were 19% slower with AI tools — yet expected a 24% speed-up beforehand and still believed AI had sped them up 20% afterward. METR stresses the result is scoped to early-2025 tools, expert developers, and familiar codebases, and does not generalize to juniors or unfamiliar code — but that caveat is the finding: the slowdown concentrates exactly where expertise and context are highest.
  • MIT NANDA (“State of AI in Business 2025”): From 150 executive interviews, 350 employee surveys, and 300 deployments, just 5% of GenAI pilots produced rapid revenue acceleration; the other 95% stalled with little measurable P&L impact. The blocker wasn’t model quality but the “learning gap” — bolting AI onto unchanged workflows. Vendor-bought projects succeeded ~67% of the time versus barely a third for internal builds, and real ROI sat in unglamorous back-office automation, not the sales-and-marketing tools that ate most budgets.

The piece names three hidden costs: a vast invisible human payroll (data workers, many in the Global South, documented by the DAIR Institute); gains spent as more work rather than less (Bloomberg’s “AI burnout era,” with one startup CEO sleeping at the office for three straight weeks — see below); and skill atrophy as engineers offload the hard parts. The “cruel symmetry”: the same economist (Brynjolfsson) whose work shows AI making junior workers more productive also runs the “Canaries in the Coal Mine” dashboard (ADP payroll data covering ~1 in 6 US workers), which finds employment for 22–25-year-olds in the most AI-exposed roles falling while rising for less-exposed peers — early evidence the entry rungs are being pulled up. The two findings are the same fact: AI most helps the novice, which is exactly why the novice’s job is cut first, optimizing away the apprenticeship that creates future experts. A field guide: AI reliably raises productivity when tasks are well-scoped, output is verifiable, the work sits inside the model’s competence, and the worker is either a novice on routine work or an expert treating AI as a draft to check; it backfires on open-ended judgment, hard-to-verify answers, experts already faster than the review loop, or orgs that buy tools without redesigning workflows.

AI anxiety and burnout in Silicon Valley. Bloomberg reported AI anxiety is fueling burnout across tech workers, with engineers “working until 4 a.m. to demonstrate productivity on par with the agents they’re deploying.” Glean’s Work AI Index 2026 (6,000 digital workers surveyed) found AI saves time but much of it goes back into cleanup (“botsitting”). A trending Hacker News essay argues AI erodes the engineering “flow state” and quietly hollows out expertise.

Ford’s AI backfire. Ford reportedly “hired AI and sacked humans,” and it backfired: the company brought back ~350 veteran quality inspectors/engineers to fix problems after its AI tools fell short, which then surged its quality-control rankings (Bloomberg). The cautionary tale is circulating precisely as the inverse of the typical pitch deck. Separately, Ford’s general counsel said in-house legal teams are adopting AI faster than many outside law firms.

Did AI kill the billable hour? The WSJ and Business Insider report consulting firms are moving away from hourly billing as AI makes work faster, cheaper, and harder to meter. Deloitte reportedly showed consultants a chart suggesting traditional labor-based consulting could shrink sharply as a share of the market by 2035, with AI agents expected to become a much larger part of professional services. Firms are testing fixed-fee and outcome-based pricing; McKinsey says more than 30% of its global fees already come from pricing tied to client outcomes. The Neuron’s analysis warns of perverse incentives: pay only for time and you reward delay; pay only for output and you reward rushed, low-quality work; and outcome-based deals that reward measurable savings can become an accelerant for AI layoffs (the easiest measurable savings is headcount). A good outcome model needs four pieces — a stable base fee, a clear target, quality guardrails, and shared upside — and should bundle metrics (cost savings + quality, speed + accuracy, revenue-per-employee + retention) to avoid hidden damage. The deeper point: billable hours and even salaries punish fast workers, and the coming fight is attribution — who gets credit when model, manager, employee, and vendor all contributed.

The reshaping of tech roles. Claude Code lead Boris Cherny posted (2M views) a theory that as engineering, product, design, and data science melt together, five role archetypes emerge: the Prototyper (invents ideas), the Builder (ships them), the Sweeper (cleans them up), the Grower (improves product-market fit), and the Maintainer (keeps mature systems secure, reliable, fast). The Neuron extends this to all functions, arguing stage-of-work (inventing, shipping, cleaning, growing, maintaining) will matter more than title or department. Kun Chen pushed back that archetypes can become “career cages,” and people should change roles as the project changes — Boris agreed; the takeaway is “know which hat the project needs right now.”

Teleoperation, robot training, and AI music labeling. Singularity Hub reported teleoperation could let companies staff “stubbornly local” jobs (e.g., operating a loader to stack potassium sulfate in a factory) with workers thousands of miles away. In India, people are being paid to wear head-mounted cameras filming everyday tasks (slicing mangoes, folding towels, making flower garlands) as training data for household/factory robots (NDTV). Tidal said it will label and restrict monetization for substantially AI-generated music. Suno launched “Spark,” an incubator offering unsigned artists grants, mentorship, and marketing — but to join, artists must let their songs be available for remixing on Suno and grant Suno a broad license to use and create derivatives of their work, with limited exclusivity and waivers around trials and class actions; the most-criticized “Good Vibes Only” clause bars artists from saying anything that portrays Suno, its staff, or products negatively, on pain of removal (The Verge, Music Business Worldwide). Kling AI (with Higgsfield AI and studio Obsidian) premiered an AI-generated WEC × Le Mans film at Cannes Lions 2026, touting motion control, multi-subject action, an industry-first 4K feature, and surpassing 100M global users and 50K enterprise clients.

Economics and labor squeeze. TLDR noted that in San Francisco even $180,000 tech salaries no longer feel like enough as housing, utilities, transport, and grocery costs surge, and a “new class of AI elite” can outspend other workers — the combined OpenAI, Anthropic, and SpaceX IPOs could mint 20+ new billionaires. EXL surveyed 322 senior leaders and found 76% of companies believe they’re ahead of competitors on AI, while by the study’s own criteria only 10% actually are — a 66-point gap (via Kamil Banc).

Hardware, chips & infrastructure

Chinese open-weight models close the gap. Meituan open-sourced LongCat-2.0 on June 30 — a 1.6-trillion-parameter mixture-of-experts model with a 1-million-token context, and the first trillion-parameter Chinese model trained entirely on domestic AI ASIC superpods for both pre-training and inference (integrating Huawei’s HCCL collective communication library for training stability). SCMP says it matches DeepSeek V4-Pro (April 2026) on benchmarks while going further than DeepSeek, which used Chinese chips only for inference — making it the clearest signal yet that China’s full-stack AI training loop now runs on home-grown hardware, landing as Washington tightens model-export controls. Zhipu AI’s GLM-5.2 ranks 6th on Artificial Analysis’ leaderboard (followed by MiniMax-M3 and DeepSeek V4 Pro) and, per WSJ/The Verge, can match leading US models at finding software bugs — cybersecurity firm Semgrep said it beat Anthropic’s Claude Opus 4.8 in some tests, and researchers said GLM-5.2 and Anthropic models can both match “Mythos” at certain bug-finding tasks. GLM-5.2 is among the most-used models on OpenRouter. The dual-use risk: hackers could use the same open-weight tools to find vulnerabilities before they’re patched, and US export limits on OpenAI/Anthropic may push more companies toward Chinese alternatives. Elon Musk said Grok 4.5 entered private beta at SpaceX and Tesla, with a 1.5T-parameter V9 model claiming performance “close to, perhaps exceeding” Claude Opus.

DRAM price-fixing class action. Fourteen consumers and three small businesses filed a class action June 25 in California federal court accusing Samsung, SK Hynix, and Micron — together ~90% of the global DRAM market — of colluding to restrict supply and lift prices roughly 700% over four years. Plaintiffs allege the coordinated pivot to high-bandwidth memory (for AI) served as cover for cutting traditional DRAM output and discontinuing DDR3/DDR4, driving up costs for everything from Apple Macs to Microsoft’s Xbox; they seek treble damages. (sedaily, Times of India.)

Super Micro / Nvidia smuggling probe escalates. Taiwan’s Keelung District Prosecutors raided Super Micro’s Taipei office on June 29 plus six residences and two affiliated companies (Chief Telecom and Albatron Technology), expanding an investigation into alleged smuggling of Nvidia-equipped Super Micro servers to China using forged shipping documents. SMCI fell ~8%. The probe already produced US DOJ charges in March against co-founder Wally Liaw, with at least 50 servers allegedly diverted in violation of US export controls. (TheNextWeb.)

Data center and national-strategy moves. Digital Realty is acquiring Blackstone’s 64% stake in three fully-leased Northern Virginia AI data centers at a $7.8B gross valuation ($3.5B cash-and-stock, 15-year hyperscaler leases, closing June 30). South Korea’s President Lee unveiled a national “Triple Axis” strategy spanning semiconductors, physical AI, and data centers (and robotics), with Samsung and SK Hynix chairs at the briefing — reported figures range from $576B (SCMP) to $880B (BBC). China’s CXMT signed a ~$3B three-to-five-year DRAM supply deal with Tencent ahead of a blockbuster STAR Market IPO. Baidu’s chip unit Kunlunxin targets a $50B Hong Kong IPO while reportedly requiring investors to buy its semiconductors — a structure BIS warned creates circular-financing risk. The BIS Annual Report warned an AI investment bubble could trigger a sharper, faster crash than a traditional banking crisis, naming non-bank debt channels as the core risk. On hardware leaks, a supposed iPhone 18 Pro/Pro Max motherboard image shows the A20 Pro chip running faster and cooler with memory and NPU upgrades for better AI performance, at roughly the A19 Pro’s footprint (AppleInsider).

Policy, safety & legal

Claude ID and biometric verification. Starting July 8, Anthropic will require accounts flagged for abuse to submit a government-issued ID plus a facial-recognition selfie, processed by third-party vendor Persona, which builds a facial-geometry template — data several US states classify as legally protected biometric information (TechCrunch). The requirement is an unusually physical demand from a chatbot.

Section 851 / US-China decoupling and the Pentagon. Section 851 of the NDAA took effect June 30, forcing DC’s biggest lobbying firms to drop Alibaba and Tencent because they must choose between Chinese tech clients and Pentagon contractors (Bloomberg). Marc Andreessen was named to a restructured Pentagon Defense Policy Board alongside Robert Lighthizer and Coleman; a16z’s defense-AI portfolio includes Anduril, Shield AI, and Skydio (Washington Examiner).

Meta’s covert chatbot probe and AI wearables privacy. Wired reported Meta hired hundreds of contractors to pose as minors while probing how competitor chatbots respond to suicide, sex, and other high-risk prompts. Internal Meta docs also show strict limits on engineers’ use of Claude Code and Codex over fears of inadvertent model distillation (The Information). Separately, Fortune highlighted the lack of legal protection around AI wearables (Omi, Plaud, Limitless AI, Meta glasses) that record meetings and can pass along likeness, voice, behavior, and geolocation for model training and third-party sale — recalling a February incident where a Los Angeles judge threatened to hold Mark Zuckerberg’s team in contempt for wearing Meta glasses into a courtroom; Meta’s AI policy states it may process information about people even if they don’t use Meta products.

Security funding. Straiker raised a $64M Series A to secure autonomous AI agents against prompt injection and tool misuse (Axios).

Opinion & essays

“The Economy of Tokens” / “RL Beyond the Verifiable.” A pair of widely-shared essays argue AI is transitioning from closed, vertically integrated systems toward a modular ecosystem built on standardized interfaces (the Transformer architecture and inference APIs). This “architectural disaggregation” lets open-weights models compete effectively with closed systems, sharply reducing costs and accelerating innovation across the stack, while stable interfaces enable specialization and independent innovation. Other circulating pieces: “It’s hard to eval is a product smell” (Hamel Husain) argues evals thinking aligns with good product design, and that with AI, verification — usually incidental during work creation — becomes the bottleneck; “Technological Involution” argues technology is approaching the upper bound of human conception as capitalism stops pressuring innovation and starts treating populations as something to be mined; and a thought experiment on a society where AI does everything, sustained by universal basic income (Fernando Borretti).

SemiAnalysis “oral history” of attention (via The Neuron). A thread traced the Transformer’s evolution: the 2017 original used Multi-Head Attention (MHA); efficiency variants followed — MQA (fewer repeated memory lookups), GQA (grouped attention, near-MHA quality at lower serving cost), and SWA (sliding-window attention over nearby text). FlashAttention was a systems breakthrough changing how GPUs store/move attention data, making long-context practical. DeepSeek’s MLA (Multi-Head Latent Attention) compresses attention info to remember context with less memory (DeepSeek-V3/R1 popularized it). For agents reading huge contexts, labs pushed linear attention and sparse attention (Gated Delta Networks, Qwen 3.5, Kimi Delta Attention, DeepSeek Sparse Attention, MiniMax Sparse Attention, GLM-5.2’s IndexShare, SWA-GQA hybrids from Cohere and Xiaomi). Serving millions of users created the KV-cache problem, addressed by Radix Attention and vLLM’s Paged Attention. Jürgen Schmidhuber noted parts of the idea trace to his 1991 ULTRA work, where “FROM/TO” played a role like today’s “key/value.”