AI Daily Digest

Friday, July 10, 2026

6,328 words · All issues

Top items

  • OpenAI launches GPT-5.6 (Sol/Terra/Luna) publicly after the US Commerce Department lifted a security-driven restriction, and rebuilds ChatGPT into a work “superapp” — introducing ChatGPT Work, merging Codex into a unified desktop app, and shutting down its Atlas browser.
  • Meta ships Muse Spark 1.1 and opens a paid Meta Model API at ~25% of rival pricing, marking a pivot from open weights toward monetization; Zuckerberg broke a three-year X silence to announce it.
  • A robotics gold rush: Agility files a $2.5B SPAC, Unitree clears its ~$5.9B Shanghai IPO, Tesla converts a Model S line for Optimus, Mistral ships single-camera navigation model Robostral Navigate, and 1X reveals a 25-DOF tendon-driven hand for Neo.
  • OpenAI retracts its endorsement of the SWE-Bench Pro coding benchmark, finding ~30% of tasks flawed — raising broader doubts about AI capability measurement.
  • Policy tightens on both sides: China weighs a “silicon curtain”/tiered open-source regime, the US operates a de facto opaque licensing regime, and a Singapore loophole lets OpenAI/Google supply models to Chinese firms’ subsidiaries.
  • News publishers seek sanctions against OpenAI for allegedly destroying evidence and misleading the court for two years about its ability to search training data and chat logs.

Company & product developments

OpenAI GPT-5.6 (Sol, Terra, Luna) goes public + ChatGPT Work + Codex/Atlas consolidation. OpenAI officially released its GPT-5.6 family across ChatGPT Work, Codex, and the OpenAI API, after weeks of delay tied to US government security concerns (officials worried the model could be misused for advanced cyberattacks against older, complex systems). The three tiers are Sol (flagship, most powerful), Terra (balanced, lower-cost everyday work), and Luna (fastest and most affordable). Sol lands slightly below Anthropic’s Fable on Artificial Analysis’s Intelligence Index but tops it on agentic coding, with upgrades to computer use, design judgment, cybersecurity, biology, and science; it was the first model to win a public ARC-AGI-3 game by correctly orienting itself in an unfamiliar scene in the game’s own vocabulary. A new “Ultra” mode coordinates multiple agents across parallel workstreams for demanding jobs, and OpenAI revealed that Sol “autonomously post-trained Luna.” Pricing matches GPT-5.5: from $5/$30 per million input/output tokens (Sol) down to $1/$6 (Luna), with Sam Altman noting “every enterprise now is thinking about spend.” Alongside the models, OpenAI launched ChatGPT Work, its answer to Anthropic’s Claude Cowork — a GPT-5.6-powered agent that can browse the web, use connected apps, edit files, operate your computer, schedule tasks, and generate slides, spreadsheets, documents, and websites, focusing on a single project for hours and pulling context to match a user’s style. It’s free on all plans on desktop and rolling out to Plus, Pro, Business, Enterprise, and Edu on web/mobile. The Codex app is being merged into a revamped unified ChatGPT desktop app with a built-in browser and computer control, and OpenAI is sunsetting its standalone Atlas browser, shifting those features into the app and a Chrome sidebar/extension. The consolidation caused notable confusion: Ethan Mollick and Simon Willison both said they were confused by the ChatGPT Work vs. Codex distinction; commentators (Theo, signüll, Corbin Braun, Daniel Lockyer) argued early Codex users wanted a focused tool, not a super-app with more modes and UI to learn, while Sriram Krishnan argued the super-app is where the market is heading. Willison documented a “dizzying matrix” of three models, multiple effort settings, prices, tool-calling, and multi-agent options. Some speculate GPT-Live/GPT-5.6 may help power the AI device OpenAI is developing with Jony Ive.

OpenAI GPT-Live voice models. OpenAI launched two new voice models — GPT-Live-1 and a smaller “mini” version — designed to sound more human and handle conversation more smoothly with full-duplex voice: they can listen and speak simultaneously, letting users interrupt naturally. Unlike the previous pipeline that stitched together separate transcription, language, and speech systems, these are unified. The mini model becomes the new default in ChatGPT voice mode, with paid subscribers getting the full version. Both can tap GPT-5.5 for reasoning and search mid-conversation, stay silent while tracking context, and surface visual information (a capability startups like Monogram are also pursuing). OpenAI executives suggested voice could become a primary interface for complex, long-running agentic tasks; it’s rolling out to free and paid tiers as rivals Apple and Amazon pursue similar upgrades.

Meta Muse Spark 1.1 + paid Meta Model API (pivot from open weights). Meta Superintelligence Labs released Muse Spark 1.1, a multimodal reasoning model built for agentic tasks, coding, computer use, tool use, and long sessions, with a 1M-token context window and the ability to split jobs across parallel subagents and operate across apps. Crucially, Meta opened a public preview of its Meta Model API — its first pay-to-use model, a pivot from Meta’s open-weight strategy amid Wall Street pressure to justify massive AI capex. Mark Zuckerberg (who broke a three-year silence on X to announce it) called pricing “aggressive and attractive” at roughly 25% of OpenAI/Anthropic rates and described the model as “state-of-the-art or very close to it” on agentic reasoning and tool use, claiming it leads Opus 4.8 and GPT-5.5 on several related benchmarks. Sources give slightly different figures: Axios/The Rundown cite $1.25/$4.25 per million input/output tokens, while The Neuron cites $0.80/$3.20 — about a quarter of top rivals either way. The API ships in an OpenAI-compatible format with structured output and parallel tool calling, and includes $20 in credits during public preview. Alexandr Wang framed coding and agentic tasks as focus areas. Meta is still training a heavier “Watermelon” successor for later this year. Separately, Meta debuted Muse Image (#3 on Arena’s text-to-image board, across Meta AI, Instagram Stories, WhatsApp) and is teasing Muse Video.

Grok 4.5 (SpaceXAI). Pitched by CEO Elon Musk as “Opus-class,” SpaceXAI released Grok 4.5 — its first model since going public — positioned as a versatile model for coding, office work, research, and writing with claimed token efficiency roughly double rival models. It shipped with Cursor as a co-developer and interest spiked in US search (EU access is delayed). Benchmarks showed it competitive with, though not quite matching, top competitors; Musk compared it specifically to Anthropic’s Opus 4.7, calling it faster and cheaper. Grok 4.5 runs $2/$6 per million input/output tokens, versus $5/$25 for Opus 4.7 and up to $5/$30 for GPT-5.6 Sol. Musk also praised Anthropic’s Mythos/Fable and promised not to “cut off” the rival. The Neuron notes parallels between Meta and SpaceXAI: both Zuckerberg and Musk tore down lackluster initial efforts before surprising with capable, cost-efficient models.

Chinese labs ramp competition. Tencent shipped Hy3, a smaller open-source model punching above its weight; Meituan fully open-sourced LongCat-2.0, billed as the first trillion-parameter coding model trained without Nvidia GPUs. Recommendations circulating for running open models to cut costs: GLM 5.2 Max for GPT-5.5/Opus level, Kimi K2.6 for Sonnet level, DeepSeek V4 Flash or Gemma 4 for Mini/Haiku level, and MiMo V2.5 Pro for GPT-5.5 Pro level. The trend: Chinese models continue to rival US labs at lower cost. MiniMax disclosed plans for a 2.7T-parameter M3 Pro model.

Anthropic tooling and governance. Anthropic introduced a “Reflections”/Reflect dashboard letting users analyze how they use Claude across 1/3/6/12-month windows, mapping habits to a “4D AI Fluency Framework,” suggesting skill creation, and adding quiet-hours nudges (beta). Claude Cowork now works across web and mobile, and access to Claude Fable 5 was extended through July 12. Anthropic also appointed former Fed Chair and Nobel laureate Ben Bernanke to its Long-Term Benefit Trust to advise on AI’s economic impacts (alongside trustees Shah, Fontaine, and Cuéllar), and invited public input on AI concerns via surveys and focus groups (“Inviting Hard Questions”).

Waymo expansion. Waymo launched fully autonomous service in Denver, Las Vegas, San Diego, and Tampa, and added the Hyundai IONIQ 5 as the first new vehicle platform for its 6th-generation Driver.

Other product/tool releases. Reve rolled out version 2.1 of its native-4K image model with element-level editing and stronger prompt adherence, retaking the No. 2 overall spot on Arena while training on under a tenth of rivals’ compute (and already outranking Meta’s just-launched Muse Image). PromptQL relaunched as “the first AI version of Slack” — a shared multiplayer AI coworker operating across a common thread on any frontier LLM, with memory that turns every correction into a reusable skill (raised $136M; relaunch video hit 2M+ views). Google Photos added a Gemini Omni-powered Video Remix tool (brighten clips, change backgrounds, add cinematic lighting, apply watercolor/sketch/oil-painting styles) for paid Google AI Plus/Pro/Ultra subscribers in the US, India, Japan, Brazil, Mexico, South Korea, Turkey. OpenKnowledge launched a local-first markdown workspace for humans and agents to co-edit knowledge bases with real-time collaboration, MCP connections, git-backed sharing, and search.

Robotics

Three humanoid companies move toward public markets in one week. Agility Robotics is going public via SPAC at a $2.5 billion valuation, merging with Michael Klein’s Churchill Capital XI for more than $620 million in proceeds. Tellingly, CEO Peggy Johnson said home robots are “10-plus years” away, so Agility is selling warehouses (GXO, Amazon, Toyota) about 1,000 Digit units on a robots-as-a-service model that has already booked over $300 million. Unitree cleared its Shanghai IPO, winning approval from China’s securities regulator to raise about 4.2 billion yuan ($618 million) on the STAR Market at a roughly $5.9 billion valuation, with a debut possible as early as late July — a rare profitable humanoid name that Chinese suppliers are watching. Tesla is converting the line that built its last Model S into an Optimus factory; Musk says output will start “quite slow” because Optimus has roughly 10,000 all-new parts with no supply chain yet, and the V3 reveal keeps slipping. AI Weekly’s framing: “the money is running ahead of the machines” — public markets price humanoids as if the hard problems are solved, while builders keep saying the near-term business is warehouses sold by the hour to customers who can fence the robot off; the home version remains a decade of hands, safety cases, and world knowledge away.

Robot brains (VLA models). Mistral shipped Robostral Navigate, an 8-billion-parameter model that lets a robot navigate with a single RGB camera — no lidar, no depth sensors — hitting 76.6% success on the R2R-CE navigation benchmark, beating the best depth-and-multi-camera systems by 4.5 points, and running on wheeled, legged, and flying robots. InternVLA-A1.5 posted the best results on all six major robot simulation benchmarks, folding language understanding and a “latent foresight” step into one policy trained on 1.2 million robot episodes — part of a cadence where a new vision-language-action model clears the bar almost weekly. LingBot-VLA 2.0 scaled robot pretraining to 60,000 hours (including 50,000 hours of real robot trajectories). GigaWorld-1 introduced WMBench, a benchmark for evaluating robot policies in learned world models. Embodied-AI lab RobbyAnt released LingBot-World 2, a world model that generates every frame in real time without a 3D engine. Nomagic’s new AI lab, headed by a former Google DeepMind researcher, claims success in early deployment of an “AI brain” for warehouse robots.

“Does VLA Even Know the Basics?” A new study tested 7 leading vision-language-action models and found they lose commonsense and world knowledge after robot fine-tuning — especially on the richer, more semantic questions their source models could answer. AI Weekly’s takeaway: locomotion is getting solved while models still lose basic world knowledge the moment you train them to act; “we’re bolting bodies onto these models faster than we’re checking what the bodies cost the mind.”

1X Neo hands. 1X detailed a new hand for its Neo home humanoid with 25 actuated degrees of freedom (22 in fingers/palm plus 3 at the wrist), close to a human hand’s 27 DOF. The tendon-driven design uses 5:1–15:1 gear ratios so every joint is natively force-controlled and fully backdrivable, with tactile sensing that reports pressure, contact location, and shear. Neo can lift a 20-pound kettlebell yet pick grapes off a stem or install a light bulb; the hands are IP68-sealed and washable, with early-access units shipping to US customers in 2026 at $20,000.

Humanoid surgery on live pigs. Humanoid robots controlled by human surgeons removed gallbladders from living pigs in a world-first experiment, using a Unitree G1 humanoid (starting price $13,500). The approach takes a fraction of an operating room’s space and is easy to deploy, potentially useful eventually for smaller hospitals and clinics lacking expensive surgical robots — though still highly experimental.

Other robotics. UBTech opened sales of its lifelike U1 companion robot for Chinese homes at around $17,650 — among the first humanoids you can actually order. Paris startup UMA, founded by a former Tesla Optimus lead, emerged from stealth with a humanoid named Northstar, signaling Europe entering the race. Nvidia is building “Halos,” a safety stack to make a 200-pound humanoid safe enough to work next to people (Neura and others are racing at the same unglamorous but critical problem). Chinese lidar maker Hesai is deepening its US business through a bigger Nvidia partnership despite sitting on the Pentagon’s blacklist — lidar becoming “the robotics version of the chip fight.” GM- and Tencent-backed Momenta had a flat Hong Kong debut after a $751 million IPO at a $9 billion valuation. Chinese analysts and outlets are openly questioning whether the humanoid boom converts viral demos into real sales. Paradigm closed a $1.2 billion fund, pushing from crypto into AI and robotics.

Research papers

Brain2Qwerty v2 (Meta + collaborators): text from brain waves. Researchers at Meta, the French CNRS, Hospital Foundation Adolphe de Rothschild, the Basque Center on Cognition, Paris Cité University, and INRIA presented an updated non-invasive system translating brain activity into text. It works in three stages: (i) an encoder — a CNN followed by a CNN/transformer hybrid “conformer” — breaks magnetoencephalography (MEG) brain activity into character embeddings and classifies characters; (ii) an “aligner” (a vanilla neural net) re-embeds and groups these by word, averaging to produce word embeddings trained to match Qwen3 embeddings of ground-truth words; (iii) a fine-tuned Qwen3-4B (with per-subject LoRA adapters averaged across subjects) corrects the sequence. Trained on 90 hours (22,000 examples) from 9 subjects typing English sentences, it achieved a 39% word error rate (vs. 43% for v1). Character error rate fell from ~50% at 20 hours of data to ~25% at 90 hours, with no plateau observed. Notably, training across multiple subjects sharply outperformed per-subject training (47.8% vs. 66.5% median word error rate) — suggesting cross-subject models improve with more data much like LLMs. Invasive electrode implants still achieve single-digit error rates, but every percent of progress reduces the need for brain surgery. Training code for both versions and v1 data are open-sourced.

DeepSeek DSpark: faster speculative decoding, open-sourced. DeepSeek and Peking University introduced DSpark, a speculative-decoding draft module that speeds DeepSeek-V4 text generation by 50%+ without sacrificing accuracy, released as DeepSeek-V4-Pro-DSpark and DeepSeek-V4-Flash-DSpark checkpoints under an MIT license on Hugging Face. In speculative decoding a small draft module proposes a block of tokens the large frozen model verifies in one pass. DSpark improves all three speed factors and uniquely adjusts verification dynamically — verifying more under light load, less under heavy load. It borrows a parallel backbone from DFlash (which uses a small diffusion model), then adds a compact “Markov head” that adjusts each position’s token probabilities based only on the previous token (fixing incoherent splices like “of problem” instead of “of course”), plus a calibrated confidence head estimating each token’s survival probability. A scheduler multiplies confidence estimates to set each request’s verification length to maximize total output across users. Results: vs. sequential drafter EAGLE-3, +30.9%/+26.7%/+30.0% accepted tokens on Qwen3-4B/8B/14B; vs. DFlash, +16.3–18.4%; gains held on Gemma4-12B. In production, DeepSeek-V4-Flash generated tokens 60–85% faster and V4-Pro 57–78% faster than the previous MTP-1 drafter. At higher guaranteed speeds, total throughput gains reached 661% and 406% (where the old drafter nearly failed) — marking new feasible operating points, converting to cheaper tokens and faster responses without touching model weights.

Google Nano Banana 2 Lite + Gemini Omni Flash. Google released Nano Banana 2 Lite (formally Gemini 3.1 Flash Lite Image), its fastest and lowest-cost image model (a replacement for the original Nano Banana), and made its video model Gemini Omni Flash available via API six weeks after consumer launch. Nano Banana 2 Lite (text/image → image up to 1K resolution) costs $0.034 per 1K image, generates in ~4 seconds, and ranks fifth on Arena.ai at 1,250 Elo — edging out the pricier Nano Banana Pro (1,245 Elo, 10 cents more) though behind GPT-Image-2 (1,386 Elo). Gemini Omni Flash (text/image/video → 720p video with synchronized audio, up to 10 seconds) costs $0.10 per second, leads video generation at 1,527 Elo, and ranks second for editing (1,347 Elo, behind ByteDance’s Seedance 2.0 at 1,377). The pair is designed to work together: a Nano Banana 2 Lite image can be passed as a starting frame to Gemini Omni Flash through the same API/conversational interface, with session history enabling up to three sequential edits. In Google’s own human-rater tests, Omni Flash took first for overall preference and instruction-following on video editing and Meta’s MovieGenBench, and tied for first on VBench I2V. Both are multimodal transformers; parameter counts and training data are undisclosed. The Batch’s take: media generation is now cheap and fast enough to run inside apps at runtime rather than as a slow production step — meeting needs of high-volume advertisers and social producers (Meta is reportedly building a system to generate ad creative including video from a product image and budget).

Anthropic + AE Studio: GRAM (Gradient-Routed Auxiliary Modules). Anthropic explored a method that could deliver the benefits of training many separately-filtered models at the cost of training only one, enabling models that inherently lack dangerous dual-use capabilities so they can be served without classifiers or filters. GRAM gives a model dedicated, removable compartments for each category of dual-use knowledge, updating only those compartments when learning from dual-use data. It resisted attackers attempting to recover removed knowledge about as well as data filtering did. Results are preliminary — not applied to production models, not tested at frontier scale. Zvi notes GRAM can’t solve inherent dual-use (the ability to “fix code” is the same as to exploit it), but works well for biological/scientific filters, and offers fine-grained account-level responses. It also raises a debate over whether “capability routing” is repressive to AI minds: Judd Rosenblatt argues routing may produce “genuine absence” and a “coherent mind whole by construction,” contrasting with RLHF suppression that leaves knowledge live in the weights with an inhibition on top (which jailbreaks exploit). Opus noted similarities to LoRA and that ablated models still showed some capability improvement.

The “J-Space”/global workspace interpretability paper (Anthropic). New research revealed Claude has a hidden “thinking space” (global workspace) where it reasons through ideas without spelling them out in its chain-of-thought — mirroring how a human might chat about one topic while thinking about another. Zvi’s extended commentary: the paper’s overall shape was expected, but the details and demonstrations were striking, providing useful new concept handles. Critics called the “workspace” definition overly broad. William Wale, who generated a notable result (a model surfacing words suggesting it thinks it’s misaligned), reports it doesn’t replicate consistently or across models. Eliezer Yudkowsky argued interpretability was once justified by claims that finding evidence of misalignment would trigger halts; Drake Thomas (Anthropic) countered that no one claimed evidence of this quality would spawn consensus for a pause, and that current models pursue problematic goals “pretty transparently with clear verbalized intent,” which he finds reassuring. Debate continued over whether the work implies Claude is conscious/a moral patient (Zvi says it doesn’t establish either, and that consciousness and moral patienthood aren’t necessarily correlated) and whether Anthropic’s anthropomorphizing is a “PR stunt” (Zvi and others argue emphatically not — Anthropic’s PR department would prefer the opposite).

Other research/technical items. NVIDIA’s Flex-Forcing trains video diffusion models to switch between bidirectional and autoregressive generation via flexible temporal and denoising-step chunking, improving inference speed, video quality, and long-range stability across compute budgets. Z.ai’s stable asynchronous RL introduced Single-rollout Asynchronous Optimization, replacing GRPO-style grouped sampling with one rollout per prompt plus value-model training and strict token-level clipping. Epoch AI published an essay (“Will we get Dyson Spheres a few years after automating AI R&D?”) arguing forecasts of speculative tech should be sensitive to how intrinsically hard the tech is to build, using a fixed drop-in-worker-on-an-H100 assumption as a lower bound — a framing Zvi critiques for ignoring recursive self-improvement (that AIs would climb the intelligence tech tree before attempting hard builds). Research on “Training-Run Assessments” (AlexM) argues scheming may be harder to detect in final checkpoints than intermediate ones, so pre-deployment evals alone are insufficient and third parties should assess intermediate checkpoints, training data, and pipeline decisions (with IP concerns the main blocker); Yo Shavit (OpenAI) called it plausibly necessary.

Policy & safety

OpenAI retracts SWE-Bench Pro endorsement. OpenAI published research finding that nearly a third (~30%) of tasks in the widely used coding benchmark SWE-Bench Pro had issues, and pulled its endorsement, calling for a more reliable option. SWE-Bench Pro was built in 2025 by Scale AI (founded by Alexandr Wang, now at Meta’s Superintelligence Labs) to supersede the “saturated” original SWE-Bench. Because such tests inform release decisions and safety criteria (e.g. OpenAI’s Preparedness Framework), flawed benchmarks distort capability pictures. UC Berkeley’s Stuart Russell told Fortune it was “very disappointing that they took this long to find out… It’s just incorrectly formulated,” adding that “our benchmarking on everything is very, very suspect” and citing dataset contamination as a huge problem. (Notably, the same benchmark underpins Anthropic’s cited orchestration/advisor cost-savings figures — see Tooling.)

US de facto licensing regime / GPT-5.6 approval muddle. Axios reported OpenAI’s public GPT-5.6 launch followed a “green light” from the Trump administration, but the White House denied this, saying no approval is required or granted and release timing rests entirely with companies — consistent with a June executive order ruling out mandatory preclearance. Yet OpenAI had previously said the administration requested it limit GPT-5.6 to government-approved customers over capability concerns. Fortune’s read: the US is operating a de facto, completely non-transparent frontier-AI licensing regime it won’t ideologically admit to. The administration has intervened aggressively before — last month the Department of Commerce forced Anthropic to pull its Fable 5 and Mythos 5 models offline over national-security concerns (later partly lifted after new safeguards), drawing criticism and sparking global “sovereign AI” panic. Departing White House AI adviser Sriram Krishnan said Trump will never establish a formal licensing regime (“a team of lawyers before you can get a model out… never going to happen”), instead favoring an ad hoc “guardrails” approach; Krishnan also reportedly supports extorting equity from major AI companies and, asked about future administrations abusing these precedents, said “I don’t think about future governments.” Trump himself spoke of “guardrails” and hinted at a forthcoming financial “contribution” from AI companies. Zvi’s “Three Pills” framing (AI/AGI/ASI) argues that influential White House figures refuse to acknowledge existing frontier-model capabilities, driving an anti-regulatory stance premised on models commoditizing — a bet Zvi and others (Dean Ball, prinz) argue is wrong if recursive self-improvement holds, since frontier labs could capture the entire Pareto frontier of intelligence, speed, and cost.

OpenAI National Security Principles. OpenAI issued principles stating it will not support: mass domestic surveillance; high-stakes decisions (including use of force) without appropriate human judgment and accountability; or uses that evade legal obligations, oversight, or accountability. It also published broader commitments (defensive cyber, biosecurity, allied government work permitted) and a plan to build “case law” through documented applications. Zvi’s critique: the principles are good in intent but hard to enforce — “mass domestic surveillance” isn’t defined in US law, “high-stakes” AI decisions already get made routinely, “appropriate human judgment” allows rubber-stamping, and the “we will not support” language is vague about what OpenAI will actually do if the government ignores it. The WSJ published new emails from the Anthropic–Department of War negotiations: Emil Michael and Dario Amodei ultimately failed to align on redlines (the Pentagon insisted on “anything lawful,” undercutting Anthropic’s autonomy and surveillance provisions with “as appropriate” and “all other applicable laws”), after which the government labeled Anthropic a supply-chain risk and raced to remove Claude from systems.

China’s “silicon curtain” / open-source regime. As US companies increasingly adopt cheaper Chinese models, Beijing is reportedly weighing curbs on overseas access to China’s top AI models. Reuters reports Alibaba, ByteDance, and Z.ai attended meetings with authorities; officials suggested leaks/thefts of AI could become punishable under national-security laws, and new restrictions on who can fund domestic AI startups are possible. A May roundtable of Chinese legal experts (per an official Supreme People’s Court journal) proposed a tiered system: basic open-source tools subject to simple filing; strategic cutting-edge open-source technologies facing export controls; and the most sensitive frontier models barred from public release or restricted to domestic use. The discussion also flagged “openwashing” as anti-competitive, data leaks, and “discourse power,” and floated using antitrust law against the open-to-closed commercial pipeline. Zvi’s analysis: both the US and China now recognize sufficiently advanced models pose national-security risks (the threshold appears to sit “somewhere between Opus 4.8 and Mythos”), and open-weight advocates demanding “Mythos-level” open models ASAP are ignoring that governments increasingly do have legitimate security concerns.

Singapore export-control loophole. The FT reports OpenAI and Google are supplying advanced AI models to Singapore-incorporated subsidiaries of Alibaba, Baidu, and Tencent — all accused by the US of ties to China’s military. Because a Singapore subsidiary is legally Singaporean, current US export controls (which target specific entities/jurisdictions, not the technology) don’t prohibit the sales. OpenAI has committed over S$300M to a Singapore applied AI lab in 2026; Google DeepMind opened a regional hub the same year; Alibaba Cloud already offers OpenAI-compatible APIs through Singapore. Commerce could collapse the arrangement overnight if it deems the sales violate the Entity List’s intent. Separately, Zvi highlights ongoing reporting that Nvidia has been “blatantly lying” to the US government, telling officials that Huawei could “satisfy global AI chip demand” — which experts say is false, damaging Nvidia’s credibility on Capitol Hill.

OpenAI copyright evidence-destruction sanctions motion. The New York Times, Daily News, Center for Investigative Reporting, The Intercept, and Ziff Davis filed a motion in Manhattan federal court seeking sanctions against OpenAI, accusing it of misleading the court about its ability to search training data and chat logs and of deleting output records. OpenAI privacy engineer Vincent Monaco reportedly revealed the company misled the court for two years about the cost/burden of searching ChatGPT logs, having actually conducted such searches before litigation began. The discovery may lead to serious sanctions in the broader copyright suit against OpenAI and Microsoft. TLDR called it a potentially “fatal misstep.”

UN AI for Good Summit & Global Dialogue (Geneva). Fortune’s Beatrice Nolan reported from Geneva. The UN’s Global Dialogue on AI was the first intergovernmental AI-governance summit bringing all 193 member states together. UN chief António Guterres appealed for worldwide AI regulation, especially on lethal autonomous weapons — noting his 2023 call for a legally binding treaty banning “killer robots” by 2026 has passed with no treaty. Conversations ranged over the AI divide between global north and south, and mitigating risks like deception and sycophancy. Creatives sought a seat at the table: ABBA’s Björn Ulvaeus argued AI wouldn’t exist without creatives. On policy, Salesforce’s Marc Benioff and Microsoft’s Brad Smith pushed back on the idea that export controls on Anthropic’s Fable 5 aimed to block foreign nationals, saying they addressed real national-security concerns — though the US actions have caused panic in Europe over losing control of a fundamental technology.

EU DSA finding against Meta. The European Commission issued preliminary findings that Meta’s Facebook and Instagram breach the Digital Services Act through “addictive design” — infinite scroll, autoplay, push notifications, and highly personalized recommendation algorithms that shift users into “autopilot mode.” If confirmed, the fine is capped at 6% of Meta’s worldwide annual turnover — over $12 billion at 2025 revenue. The Commission demands Meta disable autoplay and infinite scroll by default, introduce effective screen-time breaks, and adjust recommendations, and says teen time-management and parental controls are too hard to use to meaningfully reduce usage.

Other policy/safety. OpenAI’s national-security principles allow defensive cyber, biosecurity, and allied work while rejecting mass surveillance, autonomous weapons direction, and high-stakes automated decisions without human judgment. AI super PACs raised more than $200M to shape national AI regulation; Marc Andreessen joined a Federal Reserve AI task force (under Warsh) examining productivity, jobs, and the economy. Patreon partnered with Cloudflare to block AI training crawlers from creators’ work, with CEO Jack Conte calling for credit, compensation, and consent. The US House Homeland Security and Select China Committees are weighing federal procurement bans and contractor warnings to curb US companies’ use of Chinese AI models. Google’s SynthID watermark debunked its first high-profile deepfake — Snopes identified a fake hospital photo of Sen. McConnell after it spread on Reddit and X. FLI released its Summer 2026 AI Safety Index (grades not improving overall, Meta improved, xAI got worse). Anthropic updated its Responsible Scaling Policy to v3.4 with visible redlines (timing, redactions, confidence levels). The AI 2040 / “Plan A” project (from the AI 2027 authors, including Scott Alexander) released an optimistic policy vision built on transparency, multi-lab competition, international coordination, robot/compute fees, and a citizen’s dividend; Zvi endorses reading it without endorsing all recommendations.

Field & industry developments

Record AI venture funding. The PitchBook-NVCA Venture Monitor showed US venture capital deployment reached $412.7B in H1 2026 — nearly 30% above all of 2025 — with AI startups capturing $355.9B, or 86% of every dollar deployed. Seven mega-rounds above $1B closed in Q2 totaling $87.2B, five of them AI companies. Anthropic’s $65B round was the largest, elevating the lab to a $965B valuation and briefly overtaking OpenAI. Rounds of $100M+ now account for 87.5% of all deployment as small-check activity shrinks. Related: AI training startup Mercor is in talks for a $20B valuation, doubling from $10B in October. Databento raised a $97M Series B led by NEA to rival Bloomberg — 24 employees, profitable every month, customers include Nvidia and OpenAI. Fleek, an online marketplace connecting vintage-clothing wholesalers and retailers, raised $25M.

Temasek and MiniMax capital moves. Singapore’s Temasek will nearly triple its AI allocation to 15% over five years after its portfolio hit a record S$518B (~$400B), deploying across energy, data centers, chips, clouds, and foundation models. Chinese AI lab MiniMax launched a ~$1.9B Hong Kong fundraise: 35.6M new shares at HK$268 (a 9.9% discount) plus a HK$6.5B zero-coupon convertible bond (2.75% coupon, HK$335 conversion), to fund infrastructure and global expansion amid US chip-export pressure; the raise follows its January HK IPO at HK$165. MiniMax CEO Junjie Yan vowed to take zero salary until AGI and give up 5% of company equity, with the stock down 80% from its peak.

Data-center buildout and power constraints. Meta broke ground on its first major Canadian data center — a 1-gigawatt facility in Alberta’s Sturgeon County costing ~$9 billion over 2–3 years, its 33rd data center, chosen for abundant energy, favorable regulation, and industrial zoning. Meta is simultaneously planning a cloud business to sell excess capacity, even as investors remain skeptical of its ~$145 billion capex forecast; Meta’s stock is down ~9% this year. Meanwhile, more than $130B in US AI data-center projects were blocked or delayed in a single quarter over local pushback on power and water use. Microsoft’s 2025 carbon emissions jumped 25% to 20M tonnes as the AI buildout overwhelmed renewables, with Brad Smith conceding “sustainability solutions are not scaling fast enough.” Meta is reportedly starting production of its in-house “Iris”/”Netta” AI chip in September (designed with Broadcom, manufactured by TSMC), expecting to roughly double computing capacity to 14 GW in 2027. The Neuron/CNBC also flagged the “$3T ROI question” — whether useful AI revenue can catch up to the power plants, chips, and data centers being built.

Token pricing and commoditization debate. Multiple analyses (Benedict Evans, TLDR quick links) argue frontier models are trending toward commodity infrastructure, while Palo Alto’s CEO Arora suggested AI token costs may need to fall 90% for enterprise adoption. Zuckerberg (per Spyglass) isn’t convinced AI will commoditize, as companies already gatekeep their tech — framed as “your AI margin is Meta’s opportunity.” Zvi’s counter-view (via prinz, Dean Ball): if recursive self-improvement works, model commoditization is likely the wrong bet, since frontier labs would capture the entire Pareto frontier and tomorrow’s models may use novel architectures second-tier labs can’t quickly replicate. Analysts also note labs “moving up the stack” to own workflows, data, contracts, and habits as the model layer commoditizes, and a “bandwidth tax” spreading the memory-scarcity story into packaging, base-die logic, foundries, CXL controllers, and system integration.

Amazon CTO Werner Vogels on the “renaissance developer.” At the UN AI for Good Summit, Vogels told Fortune that AI coding tools like Claude Code (and “vibe-coding”) make reviewing and fact-checking code more important than ever — “you can’t say to the regulator, oh, AI made a mistake.” He advocates becoming a “renaissance developer”: T-shaped engineers with deep expertise plus broad cross-disciplinary curiosity (Leonardo da Vinci’s anatomy/bird-flight studies feeding his engineering). He advises engineers to spend one afternoon a week reading a paper or testing a new tool. He dismissed anxiety over AI killing entry-level jobs as “primarily noise,” saying the pace of announcements leaves even him “confused at times,” and now weighs collaboration and teamwork over raw technical fluency when hiring (open-source contributions, teamwork examples) — since programming languages can be picked up in a month or two.

Deutsche Telekom’s AI-agent network operations (case study). Deutsche Telekom handed its mobile network to a team of AI agents; response time on major events fell from an hour to a minute, and in the first month the system autonomously fixed more than a hundred problems (verified by an outside analyst). AI Adopters Club’s lesson: the AI was ordinary — success came from three “dull” decisions. Telekom aimed at one narrow job (spotting public events that flood towers, checking nearby tower capacity, adjusting before calls drop) rather than “the network.” The pattern: pick the smallest job that still matters, define “good” before building, and let the machine own only the part you can undo. Supporting data: the World Economic Forum found data quality is the number-one barrier to AI success; Deloitte found three-quarters of companies want AI agents within two years but only one in five has any real way to govern them.

OpenAI leadership shake-up. Fidji Simo, OpenAI’s CEO of Applications and No. 2 executive since May 2025, is stepping down from her full-time role to become a part-time advisor after a medical leave (a POTS relapse beginning in April) stretched longer than expected. Her exit leaves COO Brad Lightcap, CFO Sarah Friar, and CPO Kevin Weil without a defined boss just as OpenAI heads toward an IPO, with no successor named (TechCrunch flags CRO Denise Dresser as a possible expanded-remit candidate; other coverage notes her product/business responsibilities dividing among Greg Brockman, Sarah Friar, and Jason Kwon). Altman said he was “really sad about this and very grateful for all Fidji has done.”

Apple exploring larger on-device models. Apple reportedly met with startup PrismML about running much larger AI models directly on iPhones. PrismML has shrunk Alibaba’s 27-billion-parameter Qwen 3.6 to run entirely on an iPhone Pro. Larger on-device models would enable more Apple Intelligence features locally, reducing costs and enhancing privacy.

AI and the “top 1% economy” / labor. The Rundown’s Rowan Cheung, after working with headshot photographer Peter Hurley (who tells students “if you can’t beat the AIs in portrait photography, you won’t get hired”), argues the top 1–5% in service fields will get 10x more demand while the rest get replaced — because the human-services pool shrinks but concentrates at the top, and AI search recommends only one answer rather than letting people trickle down a ranked list. TLDR Founders echoed “if you want taste, you’re gonna have to eat” — taste becomes the bottleneck when AI can make thirty versions by lunch. At Brown University, a professor suspecting AI cheating switched a take-home final to in-person; 18 students dropped and 9 skipped the exam (22 of those 27 had scored perfectly on the midterm), and average scores plunged from 96 to 48 for those who took it.

Tooling & workflows

Cost-saving orchestrator/advisor patterns. Anthropic shared two patterns for keeping the expensive Fable 5 model in charge while cheaper Sonnet 5 handles token-heavy work. (1) Advisor: Sonnet executes and calls Fable only for strategic guidance or course correction; Anthropic’s advisor tool gives Fable the full conversation. On SWE-Bench Pro this reached ~92% of Fable’s score at ~63% of the cost. (2) Orchestrator: Fable plans and delegates execution to Sonnet sub-agents (“plan big, execute small”); on BrowseComp this reached 96% of Fable’s performance at 46% of cost, with each sub-agent keeping its own cache. The Rundown published a parallel guide for using ~60% fewer Fable tokens by setting Fable as planner/reviewer and routing browsing, coding, and research to Codex or a lower-cost Claude model via the Codex plugin for Claude Code or the /advisor skill. (Both note SWE-Bench Pro is now criticized/retracted.) A reader workflow (Brian H., a commercial-loan portfolio manager) uses a Claude “Commercial Mortgage Identification Prompt” to extract business mortgage filings from county recorder data, enrich contacts via web search, tier and color-code prospects in Excel, and produce an executive summary tab.

Other tools noted. Reve 2.1 (4K image editing), ChatGPT Work, Muse Spark 1.1, GPT-5.6 Sol; Runway (AI short video), RunInfra (optimize open models for production), Tabstack (extract data/automate web via one API), MixTranslate (side-by-side multi-model translations), Maxworker and Viktor/Caestro (Slack/Teams “AI employee” agents); Notion’s “Ship OS” (end-to-end product-development setup with custom agents); IBM Bob (enterprise coding agent coordinating the SDLC with governance controls); Wispr Flow (voice-to-clean-text across Claude/ChatGPT/Cursor, syntax-aware, “89% sent with zero edits”). Engineering pieces circulated on eliminating the human code-review bottleneck by delegating to agents (PostHog) and on GitHub giving 14,000+ repositories validated owners in under 45 days.