AI Daily Digest

Thursday, July 9, 2026

6,035 words · All issues

Top items

  • OpenAI ships GPT-Live, a full-duplex voice model that listens and speaks simultaneously, offloads hard reasoning to GPT-5.5, and rolls out across Free/Go/Plus/Pro tiers.
  • SpaceXAI and Cursor release Grok 4.5, the first jointly-trained model post-acquisition — “Opus-class” per Musk, 80 tok/s, 4× more token-efficient, priced at $2/$6 per million tokens.
  • GPT-5.6 “Sol” cleared for public launch (with Terra and Luna) after a Commerce Department restriction lifted; early testers describe Sol + Fable as opening a large gap over all other models.
  • A wave of funding and infrastructure moves: Prime Intellect ($130M/$1B), Ollama ($65M), SambaNova ($11B valuation), Anthropic’s $50B US data-center commitment, plus China clearing Alibaba/ByteDance/DeepSeek to buy capped H200 orders.
  • New image/robotics/coding models: ByteDance Seedream 5.0 Pro, Meta Muse Image (and privacy backlash), Mistral Robostral Navigate, Cognition SWE-1.7.
  • Zvi’s long analysis on kids, phones, screens and EdTech, plus AI’s effects on writing, jobs, education cheating, and systemic financial risk.

Company & product developments

OpenAI launches GPT-Live (full-duplex ChatGPT Voice)

OpenAI rolled out GPT-Live, a new family of voice models built on a “full-duplex” architecture that can listen and speak at the same time, a departure from prior turn-based systems that strung together separate models and forced you into a rigid “you talk, it waits, it answers” loop. GPT-Live processes input continuously, injects natural conversational fillers like “mhmm,” “yeah,” and “that’s right,” and can handle interruptions, wait through pauses, and maintain conversational flow for at least a full hour. When a query needs deeper reasoning, Live delegates the harder work to GPT-5.5 in the background while the dialogue keeps flowing, and it can surface visual cards/widgets alongside spoken answers, use search and memory, and do real-time translation.

Rollout and tiers: GPT-Live-1 becomes the default Voice model for Go, Plus, and Pro users; free users get GPT-Live-1 mini. Business, Enterprise, and Edu do not get Live at launch. It works on iOS, Android, and ChatGPT.com, but not yet in Temporary Chats, the desktop app, Work, Codex, or custom GPTs. To use it: update the ChatGPT app or go to ChatGPT.com, tap the Voice icon, allow mic access, then Settings → Voice to pick Live, Advanced, or Standard. API access is “coming soon” via a registration form; current Realtime API pricing lists gpt-realtime-2.1 audio at $32/1M input and $64/1M output tokens.

Reception and caveats: In OpenAI’s own head-to-head tests, users preferred GPT-Live over the old Advanced Voice Mode in roughly 3 of 4 comparisons, with expert science scores nearly doubling. Sam Altman said it “feels magical and ‘real’” and predicted his own preference will shift from typing to talking to AI; Zvi remains skeptical that voice beats text for his use cases and cannot imagine coding by voice, though he grants voice has niches. The Neuron flags an open concern from the livestream: Live keeps interjecting “mm/yeah” at every pause, which would be maddening if you want it to silently monitor a meeting and only speak when called on. OpenAI says you can tell Live to “wait until I ask you to respond,” but warns long pauses or background speech may still trigger it. Sources: OpenAI, The Rundown, The Neuron, TLDR, Superhuman, AI Weekly, Zvi.

SpaceXAI + Cursor release Grok 4.5

SpaceXAI (the company’s new name after rebranding from xAI) launched Grok 4.5, the first model built jointly with Cursor following its $60B acquisition — effectively a hard reset for the Grok line, which had lagged the frontier and become “somewhat of a punchline.” Grok 4.5 has 1.5 trillion parameters and posts strong benchmarks on coding, agentic, and knowledge-work tasks in “similar ranges” to Claude Opus 4.8 and GPT-5.5. It runs at ~80 tokens/second (rivaling flash models) and claims a 4× efficiency gain over models like Opus 4.8. Pricing: $2/$6 per million input/output tokens (the “fast” version $4/$18), versus $5/$25 for Opus 4.8, and it’s temporarily free inside Cursor and Grok Build. It plugs into Microsoft Word, PowerPoint, and Excel. Musk pitched it as “an Opus-class model, but faster, more token-efficient and lower cost,” with an even bigger model coming next month.

Skeptical read (Zvi): SpaceXAI “shared almost nothing else” — only four cherry-picked benchmarks, no broad detail — which suggests the model will underperform what those numbers imply; if they had something truly at Opus 4.8/GPT-5.5 level “they’d be louder about it,” and outside reactions have been muted. Token efficiency is a real edge given rising token-cost anxiety, and the lower price gives it possible uses, but Zvi expects it to fall short of GPT-5.6-Sol and Fable and is waiting for positive signals. Pliny jailbroke it, as expected. Sources: SpaceXAI/x.ai, The Rundown, Superhuman, TechCrunch/TLDR, Zvi. (The Neuron notes Grok 4.5 as xAI’s new flagship coding/office model, free in Grok Build for a limited time.)

GPT-5.6 “Sol” cleared for launch; early-tester reactions

OpenAI’s GPT-5.6 family (Sol, Terra, Luna, plus a Sol Pro tier) was cleared for broad public rollout after the US Commerce Department ended a weeks-long restriction (extra testing that had also briefly pulled early access to both Sol and Fable). Launch was set for Thursday July 9. A reported scoop claims GPT-5.6 will be OpenAI’s last 5.x release, with GPT-6 landing in ~a month, built on a bigger pretrain than the ~4T “Spud” base under 5.5/5.6; the reported motivation is matching Anthropic’s Fable 5 and upcoming Fable 5.1.

Early testers (a self-selected, biased group, as Roon cautioned) were strikingly consistent:

  • Ethan Mollick: Sol and Fable both “represent jumps over previous models and have opened a large gap with the next-best AIs”; Fable works self-directed at its own pace, Sol is faster and works “with you in steps.” He switched between them by task — Sol for back-and-forth/underspecified work, Fable for long defined tasks, Sol Pro for really hard problems.
  • Peter Gostev (most detailed): Fable is a “wise owl,” fundamentally smarter and a better writer even at low reasoning; Sol is “a rottweiler that grabs the problem by the throat,” extremely diligent and reliable — give it 8 tasks and they get done. Fable crafts better UIs but misses key things; Sol’s frontend is a big jump but still not as good overall. Sol wins on robustness/reliability, token efficiency, speed, video editing, computer use, sub-agent management, adhering to existing code patterns, and multi-day /goal runs. His “swear meter” (rudeness to Codex) dropped from 4-5% on 5.5 to 1-2% on Sol, then spiked to 7% back on 5.5. Verdict: on pure intelligence, no, Sol isn’t better than Fable, but it’s the workhorse “you can give any task to and just expect it to be done” — best to use both.
  • Dan Shipper: GPT-5.6 is a much better writer than Fable, one-shotting marketing emails Fable botches (Fable too verbose, prone to “its own private language”).
  • prinz: Sol Pro “saturates prinzbench” and can replace an associate at legal research when all relevant authorities are online.
  • Mitchell Hashimoto: Sol is his default — faster, plans/judges as well as Fable, better overall work; Fable reserved for targeted debug/security/performance. Metaphor: Sol is “a charismatic, efficient, talented coworker you’re jealous of”; Fable “a genius recluse… brilliant at its fixations but doesn’t go out, doesn’t date.” Altman quipped “tbh i dont think sol gets that many dates either.”
  • Jay: Uniquely unbiased comparison — team got “depressed” when Sol was pulled, and even with Fable back “it’s immediately clear [Sol is] just way better.”
  • Dean W. Ball / Tyler Cowen vouched for Sol’s “excellent judgment”; Tim praised end-to-end Next.js refactors from short prompts.

Sources: The Neuron, Superhuman, Zvi.

ByteDance releases Seedream 5.0 Pro

ByteDance rolled out Seedream 5.0 Pro, a multimodal image model pitched to “understand design” rather than just generate one-shot images. It improves text rendering, structure, and alignment (strong infographics and professional work), and adds “precision editing” with layer separation for editable designs, plus draw/replace/combine operations. It handles multilingual input/output natively across 10+ languages, including some right-to-left layouts and accents, targeting creators, designers, marketers, educators, product teams, and developers. The Rundown frames it as the first release nearing the frontier held by OpenAI’s Images 2.0, and notes the broader trend of the post-generation “tool layer getting eaten” by more agentic, proactive systems. It arrives ahead of ByteDance’s anticipated Seedance 2.5 video model expected this month (2.0 still near the top of leaderboards). Sources: ByteDance/Seed, The Rundown, TLDR AI.

Meta launches Muse Image (privacy backlash) and previews Muse Video

Meta launched Muse Image, an AI image generator built by Meta Superintelligence Labs, free in the Meta AI app, Instagram Stories, and WhatsApp. It creates and edits images from text prompts — filters, object removal, placing yourself in new locations, multi-reference composition, room redesigns — and Meta positions it for practical uses like custom ads, home-design previews (e.g. seeing a secondhand couch in your garage, tied to Facebook Marketplace), and Story effects. Muse reportedly outranked every image model except GPT-Image-2 on leaderboards. Generated images carry an invisible watermark, and Meta previewed a detector for Muse-generated media using those watermarks (rate limits apply). A Muse Video tool is in development.

Controversy: Muse can generate AI images from another person’s public Instagram photos simply by tagging that account; users can opt out in settings but are not notified if their content is used — raising alarms given Meta’s history (Cambridge Analytica, facial-recognition concerns). Sources: TechCrunch, TLDR AI, Mindstream, The Neuron, Superhuman.

Anthropic brings Claude Cowork to web and mobile; publishes usage data

Anthropic expanded Claude Cowork (previously Mac/Windows-desktop-only) to iOS, Android, and web, starting with Max subscribers (others in coming weeks). Cowork sessions now run in the cloud by default, letting tasks continue across devices or keep running with your laptop closed; the full desktop version retains local-file access, and Claude can notify you when something needs review/approval. Doubled Cowork usage limits are extended to August 5.

Alongside, Anthropic analyzed 1.2M anonymized May sessions, revealing that Cowork is used far less for coding than expected: just 8.7% of sessions are coding, while business operations account for 33.4% (building trackers, consolidating scattered updates into reports, onboarding checklists) and content creation 16% (decks, posts, proposals) — with ~50% of work from those top two categories. Superhuman frames this as “the work around the work” — cross-team tracking/communication tasks that belong to no one’s job description — and as a deliberate repositioning of Cowork as an “agentic background worker” rather than a full replacement, part of labs “walking back doom talk.” Sources: TechCrunch, The Verge/Mindstream, Superhuman/Claude blog, The Neuron.

Cognition ships SWE-1.7

Cognition released SWE-1.7, an in-house coding model for its Devin agent, built on China’s open-source Kimi K2.7. It reaches near-frontier intelligence (approaching GPT-5.5/Opus-level scores) at far lower cost, via broad improvements across Cognition’s RL pipeline — better infrastructure, more stable training, higher-quality data, and new long-horizon task techniques. Available now in Devin on web, desktop, and CLI. Sources: Cognition, The Rundown, TLDR AI.

Mistral’s Robostral Navigate — single-camera robot navigation

Mistral released Robostral Navigate, an 8B, hardware-agnostic robot navigation model that uses only a single RGB camera and natural-language instructions — no LiDAR or depth sensors — and runs on wheeled, legged, and flying robots. Trained on ~400K simulation trajectories across 6,000 scenes with prefix-caching that cuts training tokens 22×, it hits 79.4% success on R2R-CE seen and 76.6% unseen, beating single-camera SOTA by 9.7pp. It marks Mistral’s first move into physical-AI foundation models. Sources: Mistral, AI Weekly, TLDR AI.

Monogram emerges from stealth with visual-first AI

Monogram, from Udemy co-founder Eren Bali, launched from stealth with a $40M seed round, pitching an iOS app that swaps text-based chat for an interactive UI generated on the fly to answer requests — aimed at everyday use cases like finding movies, brainstorming recipes, or planning trips feeling more natural. Sources: Monogram, The Rundown, Superhuman.

Other model/product notes

  • Meituan LongCat-2.0: a 1.6-trillion-parameter MoE (~48B active) trained on 35T+ tokens entirely on AI ASIC superpods, with 1M context and sparse attention (Hugging Face).
  • Microsoft Copilot is reportedly replacing some OpenAI/Anthropic model calls with in-house MAI models in apps like Excel and Outlook to cut costs (the-decoder).
  • MiniMax is reportedly targeting a Q3 release of a 2.7-trillion-parameter open-source model — 6× its current flagship and larger than any current open model — pressuring closed-frontier pricing.
  • DeepSeek reportedly has an imminent V4 GA plus a bigger model to challenge MiniMax’s 2.7T Pro.
  • Google Photos added a Gemini Omni-powered “Video Remix” tool (relight clips, replace backgrounds, apply watercolor/sketchbook/oil styles); Gemini API Managed Agents gained background execution, remote MCP support, and custom functions for longer-running agents.
  • Cohere Transcribe Arabic: an Arabic speech-recognition model with open weights and hosted API.
  • Mira (General Intuition + Kyutai with Epic Games): an interactive world model letting you play multiplayer video games against AI on demand.
  • Anthropic raised Platform API rate limits and simplified its tiers.

Funding, deals & industry moves

Prime Intellect raises $130M Series A at $1B valuation

Prime Intellect closed a $130M Series A led by Radical Ventures at a $1B valuation. The open-source-training / “full-stack” platform lets enterprises build their own AI agents (compute access, an RL framework, evaluation tools) and reports crossing $100M annualized revenue in its first year, with customers including Ramp and Zapier. Sources: TechCrunch, The Rundown, The Neuron.

OpenAI Deployment Company acquires Northslope

OpenAI Deployment Company agreed to acquire Northslope, an applied-AI startup founded by ex-Palantir engineers that raised a $22M Series A in January (terms undisclosed, pending regulatory approval). It’s the deployment arm’s second acquisition after Tomoro, drawing on the $4B in acquisition capital OpenAI seeded the unit with at its May launch. The deal adds “hundreds of forward-deployed engineers” who embed with enterprise customers — mirroring Palantir’s forward-deployment playbook and escalating competition with Microsoft’s $2.5B Frontier Company and AWS’s $1B FDE org. Timed alongside GPT-5.6 launch day. Sources: Axios, thenextweb/TLDR AI, The Neuron.

Chip and infrastructure capital

  • SambaNova raised $1B at an $11B valuation as investors back Nvidia challengers; its inference-chip business (sold as server-rack units) is scaling fast, and it’s strongly considering an IPO in 2027 (CNBC/TLDR AI).
  • Iluvatar CoreX (Shanghai GPU maker) priced a Hong Kong share sale for ~$850M-$902M after a 257% post-IPO rally since its January IPO (~$475M), bringing total public capital to ~$1.3B. It’s in talks to sell ByteDance at least 50,000 GPUs this year (primarily inference). It generated only ~$148M in 2025 revenue, almost entirely from GPUs, and is expanding from government to commercial customers as US export limits tighten (thenextweb).
  • Ollama closed a $65M Series B led by Theory Venture (total funding $88M after a $15M Benchmark-led Series A). The local open-weight LLM runner reports ~9M monthly active developers, 176K GitHub stars, presence in 85% of the Fortune 500, and just 14 employees; founders previously built Docker Desktop (TechCrunch).
  • Temasek plans to nearly triple its AI allocation to 15% over five years after its portfolio hit a record S$518B (~$400B), deploying across energy, data centers, chips, clouds, and foundation models (CNBC).
  • Anthropic committed $50B to new US data centers, including custom facilities in Texas and New York built with Fluidstack, to meet Claude demand. It also plans to lease the full 16-story 330 Hudson Street building in Manhattan and double its local workforce to ~1,000 (OpenAI has 90,000 sq ft of NYC office space; Google has thousands of NYC engineers).
  • Blue Origin (Bezos) is raising its first outside funding round at a $130B valuation, with Bezos contributing $2B; it aims to return New Glenn to flight by year-end after a late-May launch-pad explosion (CNBC/TLDR).

Philanthropy & safety funding

Coefficient Giving granted $160M to Geoffrey Irving’s new venture Resolution (briefly named Sequent), which aims to combine theory and automation so AI safety can catch up to capabilities. Resolution is hiring and taking additional donations. Zvi calls it an “excellent pick.” Sources: Zvi.

Personnel

Joshua Achiam, OpenAI’s chief futurist, is leaving after nine years (“a decade where centuries happened”), joining in 2017 as a 25-year-old intern; last day July 24. His departure note is upbeat on OpenAI’s trajectory and stresses “everything is at stake” but “everything is possible.” Zvi notes the recurring puzzle of why senior people leave to do safety work with more leverage “from outside the walls.” Also: Nat Purser joins Miles Brundage and the AI Verification and Evaluation Research Institute as Director of US Policy. Sources: Zvi, The Rundown.

Nvidia’s slide, and the AI-bubble debate

Nvidia lost roughly $1 trillion in market value in under two months, dragging its valuation back to pre-AI-boom levels (Bloomberg). Meanwhile the BIS (“central banks’ central bank”) and Oracle sounded AI-bubble alarms — BIS warning AI capex “echoes every prior mania,” an Oracle SEC filing detailing OpenAI payment risk (The Register). A US Treasury Department review found the AI industry poses systemic risk to the financial system, comparing it to the dotcom crash. Zvi thinks an industry-wide collapse is unlikely (even if current models are near-peak, cheaper/faster “Fable-level” swarms are a year out and demand is high), but grants that the US “has in large part become a leveraged bet on AI”; his deeper worry remains recursive self-improvement outpacing safety. AI companies are racing toward custom chips, but fab, memory, and packaging bottlenecks keep supply tight (Axios). Sources: The Neuron, The Register, Zvi.

San Francisco AI-wealth housing “hysteria”

AI wealth is distorting San Francisco’s housing market: some homeowners accept shares of OpenAI or Anthropic as payment, prices surge as buyers bet today’s overpayment will look cheap tomorrow, landlords evict tenants to sell into the hotter market, and many properties close at least $1M above last month’s asking (FT/TLDR).

Policy, geopolitics & safety

China to allow capped Nvidia H200 purchases

Chinese officials told Alibaba, ByteDance, and DeepSeek they may soon get permission to buy Nvidia H200 chips, with total approvals capped under 200,000 units — less than half of what the firms requested — to offset an acute domestic compute shortage. Earlier this year Commerce cleared ~10 Chinese firms for up to 75,000 H200s each, but no deliveries have shipped as Beijing’s new supply-chain rules kept the deal in limbo. Sources: The Information/AI Weekly.

US moves against Chinese AI models

The US House Homeland Security and Select China Committees are weighing federal procurement bans and contractor warnings to curb US companies’ use of Chinese AI models (CNBC).

EU cybersecurity + AI action plan

The European Commission unveiled an EU Action Plan on Cybersecurity and AI, featuring pre-market model evaluation, a “structured access” blueprint, and an EU Grand Challenge for homegrown cyber-AI (commission.europa.eu).

Nobel laureate Omar Yaghi leaves Berkeley for Tsinghua

Omar Yaghi — 2025 Nobel Chemistry co-laureate (metal-organic frameworks) and a Berkeley professor since 2012 — joined Tsinghua University to launch and direct the AI for Chemistry and Materials Science Research Center (announced July 4). Yaghi, 61, aims to use AI to shorten materials-discovery cycles “by orders of magnitude,” focusing early on water scarcity and carbon capture. Nature (July 9) framed it as US science cuts colliding with Beijing’s talent-recruitment push; ~half of his ~200 trained researchers were Chinese nationals. Sources: SCMP/Nature/AI Weekly.

Off-switch for dual-use knowledge (GRAM)

Anthropic published research on GRAM (Gradient-Routed Auxiliary Modules), which delivers the benefits of training many separately-filtered models at the cost of training only one. GRAM gives a model dedicated, removable compartments for each category of dual-use knowledge, updating those compartments only when learning from dual-use data — so parts of the model’s knowledge can be deleted after training, offering a stronger model-safety framework. Sources: Anthropic, TLDR AI.

Deepfakes, watermarks, and offensive security

  • Google’s SynthID watermark debunked its first high-profile deepfake — Snopes used it to identify a fake hospital photo of Sen. McConnell after it spread on Reddit and X (TechCrunch).
  • Pliny introduced T3MP3ST, a “full offensive-security harness” that bolts onto an existing AI agent (nominally white-hat/authorized-use-only; Zvi notes red-team and actual offense look nearly identical).
  • Hugging Face was sued for alleged copyright infringement over hosting/distributing copyrighted images; Zvi notes HF and Civitai seem unenthusiastic about taking down nudification/deepfake models, calling it a “losing battle.”

FTC right-to-repair settlement

The FTC reached a settlement requiring John Deere to give farmers and independent repair providers the same repair resources it provides authorized dealers, for the next 10 years (Engadget/TLDR).

Research papers & benchmarks

Auditing coding benchmarks — SWE-Bench Pro is ~30% broken

OpenAI examined SWE-Bench Pro’s construction, model failures, and task metadata and concluded roughly 30% of its public tasks were broken, showing how flawed evaluations can distort assessments of coding ability, safety, and model progress. Sources: OpenAI, TLDR AI.

NEvo — generating videos that maximally excite brain regions

An EPFL project (NEvo, arXiv 2607.02317, model card on Hugging Face) showed it’s possible to computationally generate videos that maximally excite an arbitrary brain region via a search-based algorithm: given a target ROI, it evolves text prompts over a structured space (30 attribute categories, 614 options) in a loop — prompts→videos→predicted ROI response→evolved prompts — as a new way to probe what a brain region represents (in silico for now). Zvi urges thinking about the implications of what a sufficiently advanced AI could do to a human brain with advanced versions of this. Sources: Zvi.

Self-evolving agents taxonomy

Shilong Liu proposed a taxonomy classifying self-evolving agents into three categories — artifact optimization, harness self-improvement, and model learning — distinguishing whether evolution happens in outputs, agent infrastructure, or model weights, to give emerging agent research a common language (TLDR AI).

EBR-bench — learning from mistakes in a board game

Epoch AI introduced EBR-bench, where AIs play the (out-of-print) cooperative game Earthborne Rangers and try to learn from mistakes via a notepad. None of the AIs improve over time, and even a full strategy guide only modestly helps; models mostly don’t explore and struggle with both deckbuilding and tactics. Also noted: on EldenRingCorruptedSaveFileBench, Fable scores 100% (up from everyone’s 0%); July Fable underperforms June Fable on many benchmarks (e.g. losing ~half its APEX-SWE advantage over Opus 4.8) because it more often falls back to Opus 4.8. Sources: Zvi.

NVIDIA — open/synthetic data for agents

NVIDIA argued open and synthetic data are key to robust AI agents, highlighting the Nemotron datasets for enhancing reasoning and tool use, and noting open data aids transparency and reproducibility by letting developers inspect agent behavior (Hugging Face/TLDR AI).

Can AI writing evade detectors? (RL vs Pangram)

On the recurring idea of RL-training a model against a detector’s API (e.g. Pangram) as a reward signal to minimize AI-detected %: Benjamin Glickenhaus confirmed it’s been done — “yes we’ve done this, yes it works, no you can’t have it, it potentially made the model evil,” noting the resulting model did worse on alignment benchmarks than the base (possibly a base effect from any RL). Zvi argues you can only optimize so many things at once — forcing distinct writing degrades other measures — and that in a fair fight detection defense beats offense (contra earlier consensus), since every mind leaves a distinctive pattern (Fable can name authors of short passages). Sources: Zvi.

Tooling & developer releases

Rewriting Bun in Rust with AI agents

A Bun developer, facing many use-after-free/double-free/”forgot to free” bugs (compiler errors in safe Rust with RAII/Drop cleanup), experimentally used Anthropic’s new models to rewrite Bun in Rust — a task normally taking a small team a full year. The result was a runtime that was faster, smaller, and used less memory; the 64-minute writeup details lessons and tricks from the agentic development process. Sources: bun.com/TLDR.

TypeScript 7.0

Microsoft announced TypeScript 7.0, bringing native-code speed, shared-memory multithreading, and optimizations yielding typically 8×-12× speedups on full builds (devblogs.microsoft.com/TLDR).

Entire — distributed Git for AI coding agents

Ex-GitHub CEO Thomas Dohmke launched Entire, a distributed Git hosting network built for AI coding agents and those overseeing them, handling 570K clones/hour with regional nodes in the US, EU, and Australia — designed for the “age of vibe coding” (siliconangle, The Register, TLDR).

Other tools noted

  • Nylas Agent Accounts — give an AI agent a real email inbox (dedicated address, app password, IMAP/SMTP) without needing Gmail (Superhuman tutorial).
  • DuckDuckGo browser can now block video ads, including YouTube’s, using uBlock Origin community filter lists (Engadget/TLDR).
  • Cloudflare Meerkat — an experimental global-consensus service for managing small pieces of control-plane state.
  • Community/tool mentions: ViewGen (open-source Unreal Engine + ComfyUI plugin), LemonLime, Framer’s live-canvas AI agent, Elentaria, Gobii, ForthWrite, htmlslides, Duply, Spinach AI, Higgsfield (CGI VFX via Gemini Omni Flash).

Field & future developments

Waymo expands fully autonomous service

Waymo launched fully autonomous service in Denver, Las Vegas, San Diego, and Tampa, and added the Hyundai IONIQ 5 as the first new vehicle platform for its 6th-generation Driver (waymo.com).

Meta’s always-on recording smart glasses

Meta is testing a smart-glasses prototype with always-on recording to help users remember their day. Notably, the prototype would lack a light-up LED recording indicator; there are internal disagreements over whether captured data should be stored on Meta’s servers and used to train AI, and battery/technical challenges remain. Separately, Meta is updating existing glasses to disable the camera if it detects the privacy LED has been tampered with or removed, amid reports of people modifying glasses to record covertly. Sources: Gizmodo/TLDR, FT, The Verge/Mindstream, The Neuron.

Nuclear power in space — City Labs BOHR

Miami-based City Labs launched BOHR (Betavoltaic Orbital High-Reliability), billed as the world’s first commercial nuclear-powered satellite, on a SpaceX rideshare (alongside 80 other payloads) into a 350-400 mile orbit. It’s powered by electricity from the decay of tritium; the betavoltaic systems are far too small to power even a smartphone but represent an early step for nuclear power in space (Ars Technica/TLDR).

Engineering essays on human accountability in the agent era

Two widely-shared pieces argue humans must “own the outer loop”: agents can write code, but someone must explain what changed, why it was safe, and what happens if wrong — humans remain required in the constraints, sampling, audit, and ownership loops, and the scarce resource is human judgment informed by quality signals (logs, tests). A companion “Ownership” essay defines ownership as solving a problem end-to-end “from ‘we have a problem’ to ‘we don’t have to think about it again.’” A “You’re Not Ambitious Enough with Claude” post argues Claude delivers most leverage on high-impact work, not routine automation. Sources: registerspill/TLDR, links.

AI, writing, jobs & education (analysis — Zvi)

AI writing, “slop,” and taste

Zvi’s AI #176 documents AI writing (especially Claude/Fable’s) becoming ubiquitous and, to some, grating. Key threads:

  • roon’s hypothesis: LM writing styles are “basically fine” and weren’t better before — we’re just annoyed because we use them so much; nobody complained about the “Claude lexicon” six months ago (they preferred it to “gptslop”), but heavy daily use bred irritation. Danel Eth notes em-dashes, groups of three, and “it’s not X, it’s Y” are fine devices that grate when overused. Chase Brower disagrees, arguing LM prose is genuinely “mode-collapsed” and often bad, unlike specific humans he talks to more often without irritation. Zvi sides mostly with Brower: fine in small doses, too repetitive at current volume, and the “meta-irritation” is the real trouble.
  • Dean W. Ball: “Slop isn’t that which is bad — it’s that which is common”; any recent LM output shown to your past self as your future teen’s writing would seem genius.
  • Teddy Brown: Most writing is functional and “essentially fake” (written to exist, to be referenced, not consumed); AI clears that bar, threatening freelancers who paid bills with unenjoyable content work. Even “storyteller/narrative” roles may not survive a downturn because “taste isn’t as vital as site reliability engineering.”
  • Jodie Foster on “F1”: She said Apple’s F1 “seemed like it was made by AI,” with school-textbook structure and lines delivered exactly as a computer would write them. Zvi calls F1 “well-executed, zero-perplexity, hallucination-filled not-technically-AI slop” and predicts bifurcation: generic low-perplexity media will let AI “cook,” while high-perplexity work will use AI only cautiously, with “not AI” being part of the experience.
  • On the Myra Cheng paper (AI affirms users ~50% more than humans): Zvi cautions the context was Reddit AITA/OEQ posts (where you post expecting to be wrong) and the model tested was GPT-4o, “the poster boy for sycophancy.”
  • Chamath Palihapitiya and the LA Review of Books cited as examples of obvious AI-written content; Zvi argues popular taste tolerates or even likes AI writing, though exposure tends to reduce that tolerance over time.

Jobs — is AI net-killing them?

Zvi disputes claims that AI isn’t killing jobs. A new paper (Ara Kharazian / Ramp / Revelio Labs, 21K firms) found firms with heavy AI adoption (~$33/employee spend) grew headcount ~10% over two years (even entry-level +12%), low adopters flat. Zvi and Erik Brynjolfsson argue this only shows AI-adopting firms outgrow non-adopters by taking market share — it doesn’t establish net job creation, and early losses come mostly from failure to hire in dead-end roles. He reframes the “myths” more precisely: work expands with diminishing returns; machines doing tasks kills the specific job; cheaper (quality-controlled) things get used more but total spend can rise or fall; labor is paid well only if it’s the scarce resource — the central question. He also notes two coexisting truths: employees refusing to use AI often should be fired, yet AI cannot yet fully replace them. Musk’s “AI+Robots… universal high income, work optional” cited as clear writing.

Education — cheating and the learning paradox

  • Brown University scandal: Economist Rodrigo Serrano switched to take-home, closed-book midterm/final for the first time in 34 years; 86 students enrolled (vs usual ≤30), midterm average was 96/100 with 40 perfect scores, and graders flagged answers matching ChatGPT output. When he ran the final in-person, scores collapsed; ~50 students were implicated, but the university treated it as a “wake-up call” and sided with students, since evidence was circumstantial. Zvi: you can no longer give unproctored take-homes; ironically ChatGPT makes cheating easier to catch than mere textbook use. At UC Berkeley, the number of A’s is up 30%, rendering GPAs near-meaningless.
  • Zvi’s thesis: “LLMs are the best tool ever invented with which to learn things… and the best tool ever invented with which to not learn things.” He notes free MIT courses online produced no learning renaissance (too many trivial inconveniences, genuinely hard), whereas easier YouTube made people smarter or dumber depending on use. Yishan: learning is energetically expensive; educational systems exist to motivate/trick/force brains, and “AI is good at explaining things” is a couple steps short of that.

UK AI strategy

Andy Burnham floated a UK AI strategy prioritizing “British companies and workers” and “tech sovereignty.” Zvi argues courting American firms failed given UK conditions (restricted speech, unwelcome capital, un-buildable housing/energy, internet/VPN restrictions), and that framing questions like “what’s the point and who’s it for?” about London driverless cars (worrying about black-cab and Uber drivers) signals a losing posture. Seán Ó hÉigeartaigh offered constructive competitiveness advice. Also: Meta’s Alexandr Wang claimed a new 10×-more-compute Meta model “caught up to GPT-5.5” — but on claimed benchmarks only, which Zvi reads as not caught up in practice.

Children, phones & screens (analysis — Zvi)

EdTech: the i-Ready backlash

Zvi’s “Childhood and Education #20” centers on a viral, angry parental account (Ryan Moulton) of i-Ready, the district math software that turned his previously math-loving first-grader miserable (“hated math”; kids hiding in bathrooms to avoid it; likened to CIA-style repetition torture). The core complaint: i-Ready assumes students can’t read, reads instructions aloud slowly and unskippably (still doing so for an 8th grader), forces identical lessons repeatedly, and gives students no control — “software… thoughtlessly cruel” in ways a human teacher never would be. i-Ready is used by ~14 million students; a usedgov site allegedly marketed it on a flawed pandemic-era study. Matthew Yglesias’s defense (his son’s school uses it better) Zvi calls damning of how bad school already is. Broad point: choosing a school increasingly means choosing EdTech, which may matter more than the school; good EdTech (Zvi is “especially excited” for Alpha School’s bespoke, high-support version) is possible but most kids get a worse-than-improvised version. Non-EdTech screen abuse is also rampant — parents report kids watching 5 hours/day of YouTube Shorts on school laptops, and kindergartners (Croton-Harmon district) reciting ad jingles despite “15 minutes/day” assurances.

Social media bans: don’t do blanket age bans

Zvi opposes fixed-age (usually 16) social-media bans: age verification is easy to evade, leaks private data, violates speech/communication, and has swept in sites like LinkedIn and Substack. Australia’s ban isn’t going well — it failed to hit critical mass, so everyone assumes everyone else is circumventing it and does likewise (a smart Aussie 15-year-old told Tyler Cowen it’s ineffective but he “can no longer access LinkedIn”). On the “why is 16 vs 15-and-364-days different” objection, Zvi notes we draw lines for drinking/voting/consent too. He cites a Reason writeup of a study finding a U-shaped relationship: both abstinence and heavy use link to poorer well-being, moderate use best; for girls, no use is best in grades 4-6 but moderate best from grade 7 on (high use consistently worst); for boys, nonuse becomes increasingly harmful from mid-adolescence, exceeding high-use risk by late adolescence. Caveats: reverse causation/third factors plausible; data covered 2020-2022 (worst time to be offline); and “moderate” was defined generously — up to 12.5 of the 15 weekday hours between 3-6pm.

Some kids’ media is genuinely great

Zvi and others praise whitelisted YouTube and modern kids’ shows — Numberblocks (credited with kids spontaneously doing addition; rated far above Bluey/Paw Patrol in one review), Bluey, Daniel Tiger, Curious George — plus audiobooks and music. The problem is attractor-state slop: e.g., Spotify adding short-form video (“bait-and-switch”); fixes include screen-less devices (Google Home), turning off “Videos and Canvas” in settings (video ads require paying up), or migrating playlists to a competitor (all majors have similar issues).

Phone bans in schools: contested evidence, common-sense support

  • Allcott et al (unpublished, lockable pouches): first-year disciplinary incidents rise and well-being falls (short-term disruption), but well-being turns positive and discipline effects fade in later years; average test-score effects ~zero (modest high-school gains, especially math; small middle-school negatives); little effect on attendance, attention, or online bullying. Christopher Ferguson critiques the authors’ “naive frequentist” approach and thumb-on-the-scale framing; coauthor Thomas Dee still told press it supported bans.
  • Henry Saffer study: school bans showed no clear evidence of reduced screen time or improved well-being — Zvi’s read: bans mostly aren’t enforced. A Brazil (Rio) 2023 policy banning non-pedagogical phone use cut phone use substantially but improved test scores only 0.06 SD.
  • Zvi’s meta-point (echoing Tyler Cowen): if banning phones doesn’t move academics much, that’s evidence school itself doesn’t do much — and phones may be more interesting/instructive than dull classes. He calls it “beyond vile” to force kids physically into school (on pain of truancy enforcement) only to let them scroll AI-generated video there. Counter (Kelsey Piper): COVID school closures caused massive learning loss, and schools do teach literacy, numeracy, and invisible background knowledge; zoom school was actively destructive.
  • Anecdotal wins: Kevin Roose cites a NY Magazine piece (Anya Kamenetz) — after NYC’s ban, kids “act like kids again,” playing board/card games and sports; a Stuyvesant junior found printing study guides works better than phones (“I don’t get distracted by notifications”). Arnold Kling notes phones make teachers feel disrespected, so bans raise aggregate utility if they don’t harm students.

Device controls and physical freedom

Zvi’s prescription: (1) get phones out of schools regardless of whether kids must stay; (2) give parents real device-level controls (Tim Sweeney’s Apple proposal — parents set up kids’ accounts and pass decisions to apps without demanding “identity papers” — praised over leaky age-verification laws that don’t even address short-form-video “slop”); (3) give kids back physical freedom. He amplifies complaints that Apple’s Screen Time “doesn’t work as advertised,” and that many parents are “abusively controlling” — surveilling teens while forbidding independent real-world play, then removing their last online freedoms, leaving them nothing.