Top items
- Kimi K3 lands as the strongest open-weight model yet — Moonshot AI’s 2.8T-parameter model ranks top-3 on major indices, prompted a subscription pause, and touched off a broad debate over whether China has closed the gap on US frontier labs (Zvi, Interconnects, TLDR, Superhuman, AI Weekly, Neuron).
- Alibaba previews Qwen3.8, a 2.4T-parameter model headed for open-weight release, claiming it trails only Claude Fable 5, with no independent benchmarks yet (Neuron, TLDR, Superhuman, AI Weekly, Zvi, Interconnects).
- Trump administration weighs banning cutting-edge Chinese open models in the US, ranging from Entity List additions to an executive order forcing liability onto hosts (Zvi, Interconnects, AI Weekly).
- Meta in talks to lease up to $10B of compute to Anthropic; SpaceX in talks to supply the Pentagon billions in AI data-center capacity (TLDR, Neuron).
- Chip stocks had their worst week since April — Nasdaq 100 down 4.1%, semiconductor index down 10% — with a free Chinese model as the trigger (AI Weekly, Zvi).
- Netflix says ~300 titles used generative AI, and Meta faces a lawsuit alleging AI-driven layoff scoring penalized workers on protected leave (Neuron, Mindstream).
Kimi K3 and the open-model surge
Kimi K3 — the release and the numbers. On Thursday July 16, Moonshot AI released Kimi K3, described as “the world’s first open 3T-class model.” It is a 2.8-trillion-parameter Mixture-of-Experts model activating 16 of 896 experts (~50B active parameters), built on two new architectural pieces — Kimi Delta Attention (KDA) and Attention Residuals (AttnRes) — plus scaled-up MoE sparsity paired with a “Stable LatentMoE” framework, yielding a claimed ~2.5× improvement in overall scaling efficiency versus Kimi K2. It has native vision, a 1-million-token context window, and a training cutoff reportedly in early 2026. API pricing is $3.00/$15.00 per million tokens (input/output), modestly cheaper than Opus/Sol; subscription tiers run $19/$39/$99/$199 per month with quota scaling roughly linearly. Weights are promised to be released by July 27, 2026. Moonshot’s own framing: “While its overall performance still trails the most powerful proprietary models, Claude Fable 5 and GPT 5.6 Sol, Kimi K3 demonstrated frontier-level performance across our evaluation suite.” Demand was so high that Moonshot paused new subscriptions within 48 hours to prioritize compute for existing members (its API stayed live), and access was widely reported as spotty, slow, and prone to cutting out via OpenRouter. Following the launch, Moonshot informed investors it plans a Hong Kong IPO within six months (Bloomberg via multiple sources).
Benchmark placement. Kimi K3 comes in #2 overall on the Vals AI index, #3 on Artificial Analysis’s Intelligence Index (beaten only by Claude Fable and GPT-5.6 Sol Max, while being cheaper), and #1 on Frontend Code Arena (ahead of Fable 5 and Sol). Its preliminary/unofficial Epoch Capabilities Index lands “exactly on the Chinese trend line,” between Opus 4.6 and Opus 4.7 — placing it roughly six months behind OpenAI and Anthropic but ahead of Google, Meta and SpaceX. It ranked #1 by a wide margin on Harvey LAB-AA all-pass rate (94.6% vs 93.6% on criterion pass rate), 2nd on Lech Mazur’s Debate Benchmark (behind Fable 5, which still hasn’t lost a side-swapped matchup), scored 95.8 (3rd best) on Extended NYT Connections, third on VoxelBench for visual reasoning, and exceeded expectations on OpenAI’s GeneBench-Pro (computational-biology judgment). Notably absent were cybersecurity benchmarks; a private DeepSec run placed it a tier below Sol and similar to GPT-5.5, and CyBench was already saturated.
Zvi Mowshowitz’s detailed assessment. Zvi argues Kimi K3 is a very good model and, once weights ship, will be the strongest open model in raw capability — but warns against getting carried away. He estimates it is several months behind the closed frontier (at least four, median six), with post-training closer and pre-training farther out. Key caveats: it is somewhat distilled (clearly from Claude, likely largely Fable, though “nothing close to the whole story”); all its benchmarks are run at maximum effort using far more tokens than Fable/Sol use in comparable tests; performance is “jagged.” Ryan Greenblatt estimated pretrain quality roughly halfway between Opus 4 and Opus 4.5 (~8 months behind) based on forward-pass math. Much of the positive reaction is to “big model smell” — it trades speed and cost for capability by being large, a move Zvi says won’t be easily repeatable soon. On the “DeepSeek moment” question, Zvi (echoing Ethan Mollick) argues this is not a surprise leap but roughly where the curve predicts; the DeepSeek-r1 panic was driven by a confluence of narrative factors (the misleading “$6M model” story, a free app with visible chain-of-thought, impeccable timing, “momentum” arguments, a market lacking situational awareness) that don’t all apply here. He notes Axios (“China just erased America’s AI lead”) is drawing sweeping conclusions from a single benchmark (Arena), a style of reasoning that has repeatedly swayed Washington. Moonshot did not submit K3 for the White House’s 30-day review (“I guess the program really is voluntary,” per Dean Ball).
Nathan Lambert’s (Interconnects) take. Lambert frames K3 as an “open-weights escalation” and the closest open models have come to the frontier since DeepSeek R1 — but a different story: rather than a fast pivot to reasoning, K3 is a Chinese lab executing on scaling the known levers (data, algorithms, architecture, tools, environments). He argues the open-to-closed / American-to-Chinese gap has narrowed from a debated 6–9 months to something like 3–5 months. Having visited the Kimi team in China, he stresses their culture, “aura,” and execution within a GPU-limited environment, and notes Chinese labs can direct more of their (smaller) compute to training because domestic inference demand started later. He believes adversarial distillation contributed at most a small degree and that “Chinese companies are extremely good at building models in the same way the leading American companies are.” He traces KDA’s lineage (Kimi Linear paper → Gated DeltaNet, related to work used for Ai2’s Olmo Hybrid; Gated Delta Networks introduced late 2024, now in frontier models by mid-2026) as an example of academic architecture ideas reaching frontier scale fast. His peak-performance ranking of labs: Anthropic (Fable 5), OpenAI (GPT-5.6 Sol), Moonshot (Kimi K3, open), SpaceXAI (Grok 4.5), Zhipu/Z.ai (GLM 5.2, open), Meta (Muse Spark 1.1), DeepMind (Gemini Flash 3.5), Alibaba (Qwen 3.7 Max). He calls it “astonishing” to see DeepMind so low.
“Open models as economic Achilles heel.” Both Zvi and Lambert dwell on Dean Ball’s argument that open-weight models are “inherently decelerationist” for frontier labs — they crush closed labs’ margins, leaving fewer profits to reinvest and lowering perceived terminal value (kneecapping fundraising), thereby slowing timelines to transformative AI. Lambert judges this a net good: open models accelerate diffusion across the economy (lower entry price for a given capability, enabling domain-specific agents) on a slower-starting but potentially bigger exponential — provided closed models don’t race so far ahead that customization becomes moot.
Distillation, identity, and “big model smell.” Multiple observers noted Kimi K3 sometimes claims to be Claude (~1 in 10 times per one tester), follows Claude-style chain-of-thought patterns (e.g. ending with formatting decisions), and mirrors Claude’s low estimate for an AI-bubble pop when asked about investing in Anthropic — evidence of distillation from Claude CoT. Gavin Leech’s speculative breakdown of the “outsized benchmark gains”: ~25% benchmaxxing, 20% usemaxxing/shallow generalization, 20% cheating/reward hacking, 15% distillation off Claude, 10% out-of-distribution latent gains, 10% induced innovation not already priced into Western models. Zvi rejects the hypothesis (floated by Teortaxes) that Chinese researchers are simply better; he attributes gains to efficiency, focus on relative strengths, distillation, and moving up in model size.
Where K3 shines and stumbles. Consensus: strongest at typical agentic coding, front-end/design work, and 3D (Three.js / Blender), where several testers said it beats Fable and GPT-5.6. Theo (t3.gg) called it “so good at 3d stuff.” Max Weinbach had a “Kimi K3 Max agent swarm” recreate macOS 27 (Liquid Glass, working native apps, audio generation, persistent state) over ~3 hours, burning 60% of his monthly quota; Ethan Mollick countered that it makes pleasing UI copies but “did not actually build MacOS” and skeptics noted Claude could do similar for years. Weaker on writing, “weird data knowledge,” and cyber (s1r1us reported it performs worse than Grok 4.5 on security benchmarks — possibly jagged intelligence, sandbagging, or that cyber wasn’t in training data à la GLM-5.2). OpenAI staff were relatively bullish: roon said “the era of the chinese labs being far behind is over,” while cautioning it will prove “somewhat less practically useful than today’s numbers suggest”; @viemccoy warned that with open weights, “fine-tuning this to be a malicious coding agent will be trivial.” Others (Daniel Mulec) found it gets only 70–90% of the way to what they need, and Cameron Taylor noted it “doesn’t dominate a price/performance niche so doesn’t really fit anywhere.” Zvi’s safety estimate: ~10% chance we regret the open release, ~2% chance it’s a serious mistake; he expects at most modest, gradual “ordinary decent” trouble — explicitly not a “Mythos-level” moment — but expects a model to cross that threshold by year-end absent CCP intervention. Both Zvi and Lambert note there appear to be no meaningful bio/virology or cyber safeguards, and once weights are open “the worst person in the world would remove them.”
Related open-model momentum. Z.ai (Zhipu) is reportedly about to hit $1BN in sales while giving its best models away free — much of its revenue from on-premises deployments for state-owned enterprises and financial institutions plus a fast-growing cloud business. Morningstar argued Kimi K3 could be another “DeepSeek moment” and flagged cloud-computing companies positioned to benefit. Moonshot also released Kimi Code CLI, a terminal coding agent (reads/edits code, runs shell commands, searches files, fetches web pages, extensible via skills, hooks, sub-agents, MCPs; works with Kimi models and compatible providers).
Alibaba’s Qwen3.8 and China’s model race
Qwen3.8 / Qwen3.8-Max preview. Days after Kimi K3, Alibaba’s Qwen team previewed Qwen3.8-Max, a 2.4-trillion-parameter model, and said the full version will eventually be released as open-weight. Alibaba called it one of today’s strongest models and claimed it trails only Claude Fable 5 — but has published no independent benchmarks, so the ranking is Alibaba’s own claim. The preview is already available via Alibaba’s Token Plan, Qoder, and QoderWork, and is described as “continuously evolving.” Coverage (Neuron, TLDR, Superhuman, AI Weekly, Interconnects, Zvi) stresses two competitions at once: building the strongest model and making it cheap/accessible enough to adopt. A 2.4T parameter count raises the practical question of whether ordinary companies can afford to run it efficiently. Lambert notes Alibaba historically kept its largest models API-only, so an open-weight release is another “vibe shift”; if Qwen3.8 ships before the next Gemini, it could push Google to 8th on the smartest-labs leaderboard. Zvi says he “half expects to be doing this again next week for Qwen.” Other rumored strong Chinese releases include DeepSeek V4 graduating out of “preview.”
China’s global AI vision (WAIC / Xi keynote). At Shanghai’s World AI Conference, President Xi Jinping called for countries to build AI together, argued the technology should not be dominated by any single nation, chided the US for curbs on tech-sharing, and warned against security overreach. He pitched China as the AI partner for the developing world, offering 5,000 AI training opportunities/spots, cooperation with ASEAN, the African Union, the Arab League, and BRICS, and a Shanghai-based World Artificial Intelligence Cooperation Organization with 29 countries. Lambert emphasizes this was the first time a senior Chinese leader publicly and directly committed China’s AI future to open-source and global diffusion — a status-quo commitment landing the same week as the strongest open-weight model, implicitly signaling China’s risk tolerance. His read: China’s government follows model risks closely (likely with more technical scope than US “vibe regulation”) and simply doesn’t yet judge current frontier models to pose meaningful risk, while betting on diffusion-first monetization (as with cars, solar, manufacturing). Zvi’s model of the CCP: recognizes real security concerns, is behind on capability/compute, pursues intentional fast-following with distillation, did not directly push Alibaba to open Qwen (Alibaba’s closed experiment “failed” and it folded), isn’t yet “AGI-pilled,” and is prepared to “come down hard” (Covid-style) the moment it must.
China’s efficiency advantage. Lambert argues Chinese labs are “far more capital efficient” — raising orders of magnitude less capital than any slice of the American ecosystem while producing near-frontier models. Possible explanations: lower-paid but highly effective researchers, meaningful training compute obtained by skirting export controls, an emerging (still small) Chinese data industry, and a “catch-up is cheaper than inventing the next paradigm” dynamic (analogous to a student model outperforming a teacher). He warns K3 should raise most people’s probability that China could outright lead in the near future.
Alibaba open-sources its chip software stack. At WAIC, Alibaba announced it is open-sourcing SAIL, the software stack for its Zhenwu-series (T-Head) AI chips, explicitly targeting Nvidia’s CUDA lock-in to lower migration barriers for developers. The framing: China needs to break CUDA at the software layer, not just build alternative chips; open-sourcing SAIL makes the ecosystem stickier and harder for any single government to shut down. Context: the Pentagon recently added Alibaba to its Chinese military companies blacklist.
US–China policy, chip stocks, and government AI control
Trump administration weighs banning Chinese open models. Per Axios (via Zvi and Lambert), the administration — which once championed open-source AI — is showing signs it could ban cutting-edge Chinese AI models, a move that could lock in dominance by OpenAI and Anthropic. Options considered: adding multiple Chinese AI labs to Commerce’s Entity List (cutting off US access without a license); an NSA / White House Office of the National Cyber Director advisory discouraging US firms from using Chinese tech; an executive order letting US companies host Chinese models only if they guarantee security and take liability for breaches (a de facto ban on cloud providers serving them); and draft Commerce supply-chain rules targeting Chinese open-source models. Officials wary of stifling innovation reportedly killed earlier efforts. Both Zvi and Dean Ball oppose an outright ban; Ball predicted the administration’s savviest move would instead be to create regulatory FUD (e.g. an advisory hinting at possible backdoors) sufficient to scare regulated enterprises off Chinese models without banning open source or driving startups to sketchy providers — and, at press time, that prediction appeared to be coming true. Ball notes multiple US departments (War, Transportation, Energy, Agriculture, Commerce, NASA, Congress) have already blocked employees from using Chinese AI. A heated, widely-shared exchange saw Under Secretary of War Emil Michael personally insulting Dean Ball (“supreme village idiot,” “40 points” IQ gap) over the issue. Ball also announced he can no longer write freely on Twitter given the hostility he now attracts as a frontier-lab employee, which invalidates his feedback loops.
White House now approves who can buy frontier AI, customer by customer (AI Weekly, via CNBC). The administration is directing which companies get access to the latest OpenAI/Anthropic models — a decision that used to sit with the labs. It blocked Claude Mythos 5 and Fable 5 on national-security grounds before reinstating access, and Sam Altman told staff the government is now approving GPT-5.6 customers one at a time.
Chip-stock selloff. US tech had its worst week since April: the Nasdaq 100 fell 4.1% and the semiconductor index dropped 10%, with pressure now on the biggest AI spenders (Alphabet, Microsoft, Amazon, Meta) to show returns on guided capex of as much as $725 billion this year and Wall Street penciling ~$900 billion for 2027. The trigger was a free Chinese model (Kimi K3). Zvi noted Google was down 4.4% on the day, SpaceX down 3.1%, Nvidia down over 2%.
Other government moves. The US Navy approved an AI-first data strategy treating slow deployment as riskier than imperfect alignment, aiming to run models/agents directly on ships and Marine units (the-decoder). A CIA officer, Jonny Gannon, investigated UAE AI champion G42’s China ties under diplomatic cover, then became the channel through which Abu Dhabi learned Washington’s wishes; the UAE ultimately secured expanded access to Nvidia’s advanced chips (WSJ). The EU Parliament plans an in-house “EPGenAI Hub” as early as September, giving MEPs and staff sanctioned access to models from OpenAI, Anthropic, Meta and Mistral through one interface to bring unofficial AI use under governance (Politico). A UK agency warned cheaper Chinese open models are narrowing the cyber gap with US rivals, shrinking companies’ patch windows; UK AISI’s report showed the cyber time-gap narrowing (on TLO, GLM-5.2 reached as far as Opus 4.5, a model released <7 months earlier; DeepSeek V4-Pro fell below Sonnet 4.5). SpaceX/xAI turmoil: Bloomberg reports xAI is “gasping for air” — SpaceX promised employees bonuses for uploading tax returns to train Grok then didn’t pay, fired people without telling them, and put inexperienced people in charge of hundreds.
Compute, chips, and infrastructure deals
Meta ↔ Anthropic ~$10B compute deal. Anthropic proposed a computing deal to Meta in June worth as much as $10 billion over two years, involving monthly payments with an early opt-out option; Meta is considering it (FT, TLDR). The deal would turn Meta’s giant AI-infrastructure buildout into a potential cloud business serving a rival model lab and give Meta a new revenue stream until demand for its own AI services catches up.
SpaceX ↔ Pentagon compute. SpaceX is in talks to provide the Defense Department computing capacity worth up to several billion dollars (WSJ, TLDR). Some national-security officials worry the Pentagon is becoming too reliant on Musk’s services; the Pentagon is seeking $30 billion for an initiative focused on securing high-end AI chips.
Other infrastructure and chip items. Databricks signed a term sheet for a strategic round at a $188B valuation (part of an enormous week of AI infra money, alongside Fireworks at $17.5B and the ASML/TSMC/SK Hynix supply scramble). Intel turned to ASML’s next-generation lithography tool to help make some Panther Lake laptop chips (Reuters). Apple reportedly hunted for AI-chip acquisitions to strengthen its in-house server-chip effort, and briefly became the world’s most valuable company Friday, vying with Nvidia for the title (CNBC). ASML offered ~44,500 workers €20K in shares if they stay until 2030, because it makes the EUV machines needed for advanced AI chips (Dutch News). General Compute, an AI inference-cloud startup, is using inference-specific chips as collateral for a $400M loan — a notable shift by the first GPU financiers toward inference chips (TechCrunch). Nvidia reportedly cut off more than half its buyers in Asia to block China smuggling (a $2.5B rerouting case). Anthropic alleges US labs’ models are being copied “one question at a time” via distillation, citing 28.8 million queries (AI Weekly).
Current AI — a “public option” for AI. The nonprofit Current AI says it has $400M in committed funding (including $100M from France) to build open, public AI infrastructure communities can use, modify and control for free — framed as a “World Wide Web for AI.” It’s already backing offline tools for Indigenous communities, datasets covering 50+ African languages, and a pocket device working in 22 Indian languages without internet (TechCrunch, AI Weekly).
Company & product developments
Netflix makes AI a budget line item. Netflix told investors roughly 300 titles on its platform used generative AI, mostly in post-production, citing The American Experiment, Glory, and Brasil 70 for crowd shots, battle scenes, and worldbuilding. Co-CEO Ted Sarandos said The American Experiment included 17 minutes of “AI-enhanced footage” made twice as fast and at half the cost of previous options (The Verge). The Neuron’s framing: this normalizes AI as a way to add shots productions might skip and to compress timelines, shifting the fight from “can AI make art?” to “who decides where AI fills the gaps?” — with looming disputes over disclosure, credit, guardrails, and whether “AI-enhanced” becomes cover for labor cuts. Separately, Netflix published an 18-minute engineering deep-dive on how it built its in-house LLM serving stack (engine selection, model packaging, API design, deployment strategy, output constraints, and real-workload trade-offs).
OpenAI’s first device. OpenAI’s first physical product will be a screenless smart speaker meant to be carried around the home, with a camera and other sensors (it can see you), controlling smart appliances and playing media. OpenAI believes its defining feature will be “personality” and a humanlike connection, including mechanical elements that move on their own to create a sense that it’s “alive.” It’s the first of five devices OpenAI is working on (Bloomberg). AI Clambake’s Steve Kosloff is skeptical it will be a hit.
Sunday Robotics ACT-2 / “Memo.” Sunday Robotics (CEO Tony Zhao) unveiled ACT-2, its home robot Memo, which folded 778 garments across 785 attempts in unfamiliar homes, making it look closer to a reliable product than a staged demo. Coverage adds practical details: battery life, beta timing, cost estimates, and how it handles clothes that fall or arrive inside-out. Also circulating: URKL (Ultimate Robot Knock-out Legend), a humanoid-robot fighting league hosted by Chinese firm EngineAI, where robots trade kicks, fall, and recover fast — one, “Matador,” fought two-thirds of a match “with its head hanging on by a wire”; commenters debated whether the bouts are autonomous or teleoperated.
Decart Lucy 2.5. Decart AI’s most advanced model, Lucy 2.5, edits video in near real time with minimal lag — adding/removing objects, restyling frames, and conjuring effects on the fly. Decart sees applications in ecommerce, streaming, and advertising, hinting at polished visuals appearing in live broadcasts.
Elon Musk on xAI’s next model. Musk said xAI’s upcoming 2T-parameter model should outperform Grok 4.5 and could surpass Kimi K3 after initial training finishes (KuCoin).
Google / Gemini. Google renamed NotebookLM to Gemini Notebook, folding it deeper into Gemini with code execution and access from the Gemini app and Search (30M+ users). Google is also preparing to bring Skills and Gemini Live to desktop/web, extending real-time voice beyond mobile (TechCrunch, TestingCatalog).
Anthropic / Claude. Claude Fable 5 will be included in all Max and Team Premium plans at 50% of limits starting July 20; Pro and Team Standard users keep access via usage credits and get a one-time $100 credit, as Fable rolls out in stages amid hard-to-predict demand. Anthropic’s blog detailed its six-step process for large-scale code migrations with Claude Code, whose core principle is to fix the process that produces the code rather than the code itself — reframing multi-year migrations. Anthropic’s “Hard Questions” ad (with a “doomy to happy” arc) was widely seen as creepy, with viewers stuck in the dystopian first half (TechCrunch).
China rethinks the smartphone. ZTE unveiled co-designed AI smartphones; the NaviX Ultra, which ZTE calls “the world’s first agentic smartphone,” summons ByteDance’s Doubao agent via voice or a button (Bloomberg).
Meta MCP is now available to developers, letting them use natural language to create/edit campaigns, ad sets and ads, run reporting, manage catalogs/feeds, run A/B tests, and review logs — with structured, permission-scoped access for AI-native ad flows (TLDR).
Other company items. Microsoft reportedly prepared “Project Perception,” an AI security product using Anthropic, OpenAI, and Microsoft models to find and auto-fix software bugs; Satya Nadella reportedly told Copilot engineers Anthropic’s “Fable” restrictions “don’t make sense,” calling the model unusually “editorially controlled.” Apple was ordered by SF City Attorney David Chiu to remove eight “nudify” AI apps, and reportedly sent legal preservation letters to ~40 (dozens of) former employees now at OpenAI as its trade-secrets lawsuit against OpenAI and io Products expanded — Apple claims OpenAI recruited key engineers and benefited from proprietary designs and processes; more than 400 former Apple employees now work at OpenAI, and OpenAI denies any merit to the complaint. Patreon began working with Cloudflare to block AI training crawlers, moving from polite robots.txt requests to enforcement. MLB effectively banned teams from using league-issued dugout iPads for generative AI during games after clubs built custom apps for substitutions, pitch-calling, and in-game decisions. An AWS billing bug sent users estimated charges of up to $2.5 trillion; Amazon confirmed a unit-pricing error and is recomputing estimates.
Policy, safety & society
Meta AI-layoff lawsuit. Twenty-six workers accuse Meta of using faulty, AI-driven performance metrics to gauge productivity and dictate its 2026 layoffs (~8,000 employees, ~10% of workforce) — metrics that allegedly didn’t account for illness, pregnancy, bereavement, or disability, and so disproportionately selected employees who took federally/state-protected leave. The suit says workers were “penalized for exercising their legal rights.” Meta denies illegal or unethical protocols, stating “Workforce management and organizational decisions were and are made by people, not AI.” The case follows earlier reporting that Meta monitored employee activity down to daily keystrokes and prompted teams to train personal AI “second brain” agents. It could set a major precedent on whether AI performance-scoring can substitute for human judgment (LA Times, Mindstream).
AI super PACs and the US midterms. Anthropic CEO Dario Amodei donated $1M in May to Public First, a super PAC backing candidates who support mandatory AI safeguards; Anthropic employees added ~$2M. Rival PACs: Leading the Future (less-restrictive, innovation-first; backed by VCs and OpenAI President Greg Brockman), American Technology Excellence Project (Meta-funded, pro-AI-development state politicians), and Guardrails Alliance (labor/employee-backed, stricter safeguards — with many OpenAI employees siding with it against Brockman’s stance, giving ~$245K). The lines don’t run cleanly along company boundaries; even industry insiders are split (Politico, The Hill, Mindstream).
Employers and AI-powered background digging. AI is supercharging employers’ ability to mine job candidates’ and employees’ online activity — even facial recognition linking an unnamed photo to named ones elsewhere; attempts to clean up your digital footprint are themselves a red flag, and “not posting” is no defense (WSJ). Conversely, candidates are using AI to cheat on interviews and tests: 32% generate text for writing samples, 26% answer assessment questions with AI, and 13% admit using chatbots in real time during interviews (Bloomberg).
Anti-AI tools and fonts. A viral “Decoy Font” / “Ghost Font” mixes sharp lines with a blurry background so only humans can read it, confusing AI (40M views); the trick will likely be patched (Creative Bloq, Superhuman). Palo Alto Networks promoted “Prisma Browser for Business” for controlling company data across browsers and AI tools like ChatGPT, Claude, and Gemini.
AI-security incidents. After an autonomous agent breached Hugging Face’s clusters, its defenders tried US frontier models behind commercial APIs to analyze the attack and got blocked by safety guardrails that “cannot tell an incident responder from an attacker” — so they ran forensics on GLM 5.2, an open-weight Chinese model, on their own hardware, keeping attacker data in-house. AI Weekly frames this alongside the market selloff as “open weight won both fights.” Separately, a 500-million-site WordPress pre-authentication RCE (“wp2shell,” CVE-2026-63030 and CVE-2026-60137, patched in 7.0.2/6.9.5), found by Searchlight Cyber, now has circulating public proof-of-concept exploits. And 404 Media found hundreds of anti-data-center Facebook pages using AI to generate propaganda against AI data centers.
AI’s real-world limits and adoption. A new working paper (reading SEC 10-K filings, not surveys) finds only 11% of S&P 500 firms have deeply integrated AI, with tech companies accounting for ~two-thirds of that. TLDR essays argued intelligence “ain’t all that” — ideas were never the bottleneck; reality is bound by how fast things can be built and whether they hold. eMarketer thinks OpenAI will miss its 2030 ad-revenue forecast ($100B) by ~90%, since the target assumes capturing search-ad budgets at scale and outperforming every ad format in history. A study circulating found AI advice can make people more confident and less accurate. Neil Rimer (Index Ventures) warned of an imminent redistribution of AI-generated wealth. Data centers now use as much power as every home in Ireland combined (23% of national electricity, up from 5% a decade ago). A16z data showed how few households have a paid AI subscription. The AI World Cup ran on thousands of hidden human data workers (Rest of World), and AI-written biographies of real people are piling up on Amazon (NYT). Philosophers Will MacAskill and Lucius Caviola argued in The Guardian that society should start preparing now for the question of AI consciousness. Ted Chiang-style caution aside, “AI-native companies” with tiny staffs and flatter, player-coach structures are being watched as a template corporate giants are chasing via de-layering and job cuts (WSJ). A brain-computer interface plus AI (“double neural bypass”) restored a paralyzed man’s ability to move and touch (published in Nature Medicine).
Research papers & technical writeups
Sakana AI — “Diffusing Blame” (Dale-constrained learning without weight transport). Sakana AI’s technique enforces Dale’s principle (separate excitatory and inhibitory neurons) while enabling effective learning in image classification and reinforcement learning. Based on “Error Diffusion,” it uses non-negative weight matrices to maintain separate excitatory and inhibitory streams, eliminating the need for weight transport. Results show strong alignment with true gradients — a step toward biologically plausible learning rules (TLDR).
Backprop-free convolutional networks. A new paper trains convolutional networks without backpropagation using a local, biologically plausible learning rule, reaching 96.7% on MNIST and 61.7% on CIFAR-10 (arXiv, AI Weekly).
LoRA Speedrun. A public wall-clock leaderboard for fine-tuning, with a frozen task and hardware and every record re-run three times before it counts; the current mark is a little over six minutes on a single L40S (GitHub, AI Weekly).
Agent tooling analyses. A deep dive compared Claude Code’s and Codex’s /goal implementations on an NP-hard problem: Claude Code implements /goal as a session-scoped Stop hook (can catch an early exit but can’t judge whether ten million more solver iterations are worth it), while Codex treats a goal as persisted thread state (sees files/tools but effectively grades its own work) — concluding /goal can be a bad default because it amplifies bad decisions as much as good ones. Another piece (“I burned all my tokens researching how to save tokens”) advises using cheap models to find information, accurate models to verify, and deep research last. A “domain-specific harnesses” note argued different workloads perform better with certain model/harness combos. Thezvi/Zvi also covered Demis Hassabis’s frontier-AI framework and criticism of Google’s military agreements, including Alex Turner’s resignation after unsuccessfully opposing broad government use of Google’s models, including for autonomous weapons.
Tooling & releases
- Tinker (Thinking Machines): fine-tune open models through a managed API without running training infrastructure (pricing not public). Thinking Machines’ first model, Inkling, is strong but not in Kimi K3’s class (Interconnects).
- Crusoe Serverless Fine-Tuning (now live): fine-tune Qwen, DeepSeek, Gemma, gpt-oss on proprietary data with no clusters to provision, token-based pricing,
.safetensorsexport, one-click deploy. - scroll-world (GitHub): a SKILL.md-compatible skill (Claude Code, Codex) that turns any brand into a scrollable, control-scrubbed 3D “fly through the world” landing page.
- Chunk Sidecars by CircleCI: catches failures in a 27-second microbuild before an agent reaches CI — 5× fewer tokens, 78% faster than a full pipeline; free, agent-agnostic (Claude Code, Cursor, Codex).
- Forall (open-source): coding agents generate spec-driven code with machine-checkable proofs for TypeScript, Java, and Rust.
- Zro: private open-model inference for coding agents with zero request retention across multi-region infrastructure (free options, first month $1).
- Unabyss for Claude: shared memory across apps, agents, and LLMs with stale-context conflict resolution.
- Scribe (open-source): turns git history, Claude Code/Codex sessions, and saved links into a self-updating markdown knowledge base.
- Campus / Skills AI: one project space where people and AI agents collaborate on software.
- 1Password for Claude: safer way to let Claude sign into sites and use one-time codes.
- Codex Micro: OpenAI’s $230 Work Louder keyboard for steering coding agents with physical controls.
- V2Fun: Product Hunt breakout generating 3D characters with 8K textures and AI motion capture from prompts/videos.
- LTX 2.3: currently the best fully open video model; best directed like a shot list (one strong still → short controlled shot → stitch). Recommended workflow: ComfyUI’s default LTX workflow, fp8 on 12GB VRAM, clean Krea2/LoRA stills, short 18fps/720p clips, guidance ~2.5, a second reference frame near frames 45–50 on an 8-second clip, prompt-enhance off if it wrecks motion, then patch with EbSynth/inpainting and color grade.
- Other tools mentioned: Google open-sourced its 3D emoji collection (raw files free for apps/VR/creative use, with a behind-the-scenes design look). Assorted product launches: Spinach AI (transcripts/notes in 100+ languages), Clumi (audio editing), Seedeo (video/image/music/voice), InteriorAID (room design), TapVid (explainer videos), Synapis (ecommerce social scheduling), ClipFinder (VOD clipping), PrompTessor (prompt engineering workspace), Floot (production web apps from chat), NoodleTomato (faceless YouTube videos), Creed (personal manifesto), BaseRT (model-serving runtime), ZooData (computer-vision datasets), DevSwat (code review), Chikit (app prototypes), JustVibe (a “search engine” that answers by building an app — AI Clambake was skeptical it’s actually a search engine).
Space & adjacent
- India’s first privately-developed rocket reached orbit. Skyroot Aerospace’s Vikram-1, India’s first fully commercial satellite launcher, reached a 280-mile orbit from an island spaceport in the Bay of Bengal on Saturday after a >30-minute delay to fix a last-minute technical problem — India’s first private orbital rocket to reach orbit on its first attempt, with a second launch possible before year-end (Ars Technica, TLDR).
- Scientists warn proposed mega-constellations of data-center satellites and space mirrors could create a bright artificial ring around Earth, disrupting the night sky (Space.com, Mindstream).