Top AI News Weekly

Week 36 · Aug 31 – Sep 6, 2026

Six Labs in Seven Days + H3 Became Everyone's Substrate. OpenAI previewed GPT-6 Astra, Google shipped Gemini 3.8 Flash, Anthropic released Claude Fable 5.1, DeepSeek opened a 305B multimodal V4, Alibaba served Qwen3.8-Max-0902 on a 1M context and Meta updated Muse Spark - inside one week. The pattern underneath is who is allowed to use them: Gemini 3.8 Flash Cyber goes only to vetted defenders through Google's Fairwind Program, and Claude Mythos 5.1 is Fable with lower safeguards, restricted to verified cybersecurity professionals and life scientists. Deliberately less-guarded models, gated by profession rather than by payment, is now a shipping pattern rather than a thought experiment. The week's other story is MiniMax's H3. Nine projects here are built on it: VDN's hybrid attention claims 2.65x per layer, a ComfyUI turbo pack gets it down to four steps, Viggle-Animate finetunes it for character replacement at 26 seconds a clip needing no pose skeleton or mask, two unrelated groups both shipped an interactive world model called H3-World, and SolarWM uses it as one of four backbones. Open weights stopped being a distribution choice and became infrastructure other people build businesses on. Against all that, the most quietly impressive result cost 67 cents: a 75M-parameter transformer trained from scratch reaching 44% on ARC-AGI-1, no LLM involved. Google mapped a complete male fruit-fly nervous system at 166,000 neurons and 125 million synapses, and OpenClaw shipped 2.0 out of 16,000 pull requests from 933 contributors.

64 launches and research drops that matter for enterprise AI builders—curated, tagged, and ready for your next roadmap sync.

New drops

64

Unique sources

58

Key themes

Developer · Agents · Video

language models

Language Models

LLMs, reasoning models, MoE architectures, and language understanding systems.

Claude Fable 5.1Anthropic

Claude Fable 5.1 and Mythos 5.1

Fable 5.1 reaches general availability on AWS, Google Cloud and Azure, with Anthropic citing 52.6% on Terminal-Bench-Science against Fable 5's 24.7%, 55.8% on Terminal-Bench 4.0, and parity with Fable 5 at lower effort settings. Pricing is $10 per million input and $50 per million output tokens with a 75% cache-read discount, roughly 25% cheaper overall for typical workloads. Mythos 5.1 is the same model with lower safeguards, released only to verified cybersecurity professionals and life scientists through trusted access programmes - the second profession-gated release this week.

View release ↗
DeepSeek V4 VisionDeepSeek

DeepSeek-V4-Flash-Vision-Exp

DeepSeek's first multimodal model in the V4 family: a 305B mixture-of-experts built on the V4-Flash architecture with visual modules added, plus DFlash attention, hyper-connections and DSpark speculative decoding. Released under MIT in BF16, F32, F8_E4M3, I8 and I64, running on vLLM and SGLang. DeepSeek reports 36.5 against 26.2 on ApexBench for multimodal agent work over V4-Flash-0731, while holding text-only roughly steady at 83.9 against 82.7 on Terminal-Bench 2.1. Context length is not stated on the card.

View model ↗
Qwen3.8 MaxAlibaba Qwen

Qwen3.8-Max-0902

Served through Qwen Cloud with image, text and video in and text out, on a 1M-token window - up to 991K input and 131K output, or 983K in thinking mode with a 262K reasoning-token cap. Pricing is $2 and $6 per million standard input and output tokens, with cached input from $0.17 on explicit reads, at 1M TPM and 15K RPM. Qwen claims better coding, collaborative agent performance and vision understanding than prior versions. No parameter count, architecture, licence or benchmark score is published - those fields are genuinely absent from the page rather than omitted here.

View release ↗
Gemini 3.8 FlashGoogle

Gemini 3.8 Flash and Flash Cyber

Announced together on 2 September. Flash scores 54.9% on HLE-Verified, leads Vals Finance Agent V2 and Harvey's Legal Agent Benchmark and beats larger frontier models on DeepSWE v1.1 - all vendor-reported - at an introductory $0.75 and $3.75 per million tokens, rising to $1.50 and $7.50 after 31 December. Flash Cyber is the restricted sibling built for finding and patching vulnerabilities: past 70% on an internal 20-language vulnerability benchmark, 47.2% pass@1 on CWE-Bench, and per Chrome Security 2.6x more correct patches than the best commercial models. It carries deliberately more permissive cyber-defence mitigations and goes only to vetted defenders through the Fairwind Program.

View release ↗
GPT-6 AstraOpenAI

GPT-6 Astra

Released 3 September as a limited preview for trusted partners, with OpenAI claiming state of the art on computer use, browser use, software engineering, cybersecurity, science and professional work, and better adherence to task boundaries than its predecessors. It says the model earned perfect or near-perfect scores on key reasoning benchmarks, ahead of GPT-5.6 Sol and Claude Fable 5, though no specific figures were published. Rollout starts with cybersecurity partners before reaching paid ChatGPT tiers and the API. Note: openai.com returned 403 to our fetch, so this entry rests on secondary coverage rather than the announcement itself.

View release ↗
Agentic Coding ModelMeta AI Research

Muse Spark 1.3

Released 2 September through Muse Code on macOS and Linux and the Meta Model API, aimed at agentic and coding work. Meta reports roughly 20% fewer tool calls and 25% fewer tokens on coding tasks than Muse Spark 1.2, plus better long-horizon execution, single-thread multitasking, self-awareness of its own limits and prompt-injection robustness. A scorecard compares it with 1.2, GPT 5.6 Sol and Opus 5 across agent, coding, instruction-following and long-context evaluations, but the figures are published as an image rather than text. Closed and API-only for now, with open weights stated as on the roadmap.

View release ↗
vision & image

Vision & Image

Image generation, editing, restoration, OCR, and visual understanding.

Diffusion Image ModelinclusionAI

LLaDA-Image

A 6B diffusion text-to-image and instruction-editing model where both the backbone and the DiT are diffusion models trained in one unified framework, with VQ-conditioned generation and Chinese and English text rendering. Samples in 50 steps, or four for the distilled Turbo variant. Weights open on Hugging Face and ModelScope in BF16 and FP8. inclusionAI reports state-of-the-art overall scores of 53.53 English and 53.38 Chinese on Qwen-Image-Bench, without naming an external baseline. A licence file is referenced but its terms are not stated on the card.

View repo ↗
Omni-Visual Diffusion LLMShanghai AI Laboratory

InternLumina-U2

A multi-codebook diffusion LLM unifying text, image, video and 3D understanding with image generation and editing in a single MoE diffusion backbone, working on discrete VQ-level visual tokens with parallel spatial-position and sequential-codebook prediction. Reported against InternVL-U, LLaDA2.0-Uni and LLaDA-o: CharXiv-DQ 83.65 against 69.80, HallusionBench 62.15 against 50.20, MathVision 33.22 against 26.70 and DynaMath 56.37 against 46.97. Inference code is up, but the technical report and weights are both still marked coming soon, and no parameter count or licence is given.

View project ↗
video & animation

Video & Animation

Video generation, editing, inpainting, and real-time rendering.

Faceless Shorts PipelineHasan Aboul Hasan

Claude Faceless Shorts Creator

A Claude Code system for faceless YouTube Shorts across three pipelines: pure-code Remotion compositions in TSX, AI-generated video through fal.ai, and documentary-style collages. ElevenLabs handles narration and drives word-exact synced captions. It ships a beat grammar - hook, setup, quiz, reveal, twist, loop - plus a reusable brand contract file and twelve worked examples, with Python utilities over ffmpeg, PIL, rembg and Playwright. MIT, 218 stars.

View repo ↗
Agentic Video UnderstandingGoogle

Gemini Agentic Video Understanding

Gemini 3.8, 3.7 and 3.6 Flash and 3.5 Flash Lite gain an agentic video mode loading frames on demand instead of sampling at a fixed rate, cutting token use by up to 88% on long content against static sampling. Models with the 1M-token window take video up to three hours at low media resolution or one hour at high. Default static sampling stays at one frame per second, costing roughly 100 tokens per second of video at low resolution or 300 at high, and public YouTube URLs can now be passed directly in a free preview.

View release ↗
YouTube Clip ExtractorAIONIX

OpenClip

A local-first tool pulling a YouTube video and cutting short vertical clips for TikTok, Reels and Shorts: an LLM picks the moments, MediaPipe tracks faces and reframes to 9:16, active speaker detection runs off the audio, and Whisper writes the captions. FastAPI behind a Next.js front end with SQLite project storage, MIT, 7 stars. Despite the name it has nothing to do with OpenAI's CLIP or the open_clip library.

View repo ↗
Action-Conditioned World Modelhugging-apps

H3-World Action Demo

An action-conditioned world model on a rank-32 LoRA finetune of MiniMax-H3: give it a first frame and it responds to WASD movement, J and L pan, I and K tilt and an F sprint modifier. The conditioning is the clever part - it rides entirely on text, one short English sentence per latent frame appended to the scene prompt, held in place by a directed attention mask so each caption binds only to its own frame. Because H3 is 195.9 GiB in bfloat16 the pipeline splits across two Spaces. Defaults to 960x544, 28-step or 8-step turbo sampling, trained on Apache-2.0 ABot-World-Explorer-500h footage.

View model ↗
Real-Time AI Streamingblendi-remade

Unreel

A demo streaming service writing and rendering each episode shot by shot while you watch, so nothing on it exists until you press play. Ships 14 continuous-shot films across sci-fi, noir, western and fantasy plus eight hard-cut comedy loops, using fal's MiniMax H3 Max Turbo for video, Nano Banana 2 for cover art and Gemini 2.5 Flash to write the episodes, chaining image-to-video for frame continuity. Next.js 15 and React 19, cited at roughly under a dollar per ten-minute session, 126 stars, no licence stated.

View repo ↗
H3 Turbo NodeLarryvrh

ComfyUI-MiniMax-H3-Turbo

A ComfyUI node pack - a turbo LoRA loader and a matching sampler - getting MiniMax-H3 to produce synchronised video and audio in four to eight sampling steps instead of about twenty, for both text-to-video and image-to-video. Takes full, pruned or quantised base weights, wants width and height as multiples of 32 and 124 to 362 frames, has a low_vram toggle trading sharpness for memory, and detects the ComfyUI version to schedule audio correctly. Ships a v4 checkpoint plus a v1 that holds up better on four-step heavy motion. Apache-2.0, 553 stars.

View repo ↗
Interactive World ModelTencent, NUS and Hong Kong PolyU

H3-World

Turns the 33B MiniMax-H3 into an interactive world model by converting its existing zero-shot natural-language control of characters and camera into precise, temporally grounded control, without bolting on a dedicated action module. Each action is a structured character-and-camera instruction aligned to video latents, with temporal attention routing confining an instruction to its intended interval so control does not leak across actions. Trained on 8,000 gameplay samples over 10,000 LoRA steps touching 0.199% of parameters, and it still generalises to unseen scenarios. arXiv paper dated 1 September; unrelated to the community Space sharing the name.

View project ↗
Long-Horizon World ModelCUHK-Shenzhen, NUS, HKUST, NVIDIA and MSRA

SolarWM

An open video world-model framework trained on five-second clips that runs minute- to hour-scale real-time inference while keeping interaction responsive and the world coherent. Four backbones at different scales - SolarWM-Wan-5B, Wan-14B, LTX-2.5 and Minimax-H3 - across causal student and bidirectional models. Ships an open dataset of 1.43M clips at 25TB with the full processing pipeline including camera annotation and filtering, plus code and checkpoints; dataset access needs a request form.

View project ↗
Character ReplacementViggle AI

Viggle-Animate

A full finetune of MiniMax-H3's ref2va transformer at 33.1B parameters, jointly distilled with DMD2 down to three forward passes plus a rank-128 LoRA delta. What makes it notable is what it does not need: no pose skeleton, segmentation mask, face crop, background plate, depth map or text prompt - just a driving video and one repainted reference frame. Renders 124 frames at 480x832 and 24fps in 26 seconds on a single B200, which Viggle puts at 6.1x faster per clip than Wan2.2-Animate-14B on equivalent hardware. Weights open under the MiniMax H3 community licence, inference code Apache-2.0.

View model ↗
Interactive World ModelRunway

GWM Worlds 2

An autoregressive diffusion model for real-time interactive world simulation, producing continuous 720p 24fps video with generated 48kHz audio that responds to text actions and camera control through causal KV caching. New over the first GWM Worlds: audio, richer subject and scene control, and multiplayer with independent per-subject control, running indefinitely rather than to a preset duration. Labelled a research preview on Runway's platform only, with no parameter count, no named-baseline benchmarks and no open release.

View release ↗
audio & speech

Audio & Speech

TTS, STT, voice cloning, music generation, and audio processing.

Breeze TTS 2BreezeBlue

Breeze TTS 2

Released 7 August as the successor to BreezeBlue's first TTS, aimed at interactive voice work in games, animation and voice agents. Voices are designed in natural language by describing character traits rather than picked from a list, and directed inline with tags for emotion and pacing; it streams over WebSocket across roughly 50 languages with accent control. On BreezeBlue's own benchmarks it leads Voice Design at 78.02 Role Fit, 5.24 points clear of the runner-up, Voice Direction at 4.25, and latency at 119.4ms median time-to-first-byte - all vendor-run, not independent.

View release ↗
Realtime TTSInworld AI

Realtime TTS-2

Conditions on multi-turn audio history rather than a single sentence, inferring emotional state, tone and pacing from how the conversation has actually gone, and replaces fixed emotion enums with natural-language direction. One voice identity holds across more than 200 languages, with language switching inside a single generation. Inworld states sub-100ms time-to-first-byte and a first-place Artificial Analysis ranking for realtime TTS from blind testing. Proprietary and hosted - no weights and no model size disclosed - through Inworld's API and via Cloudflare and DeepInfra.

View release ↗
Tokenizer-Free TTSOpenBMB

VoxCPM2

Tokenizer-free TTS: an end-to-end diffusion-autoregressive architecture generating continuous speech representations directly rather than predicting discrete audio tokens. 2B parameters trained on over 2 million hours, covering 30 languages plus nine Chinese dialects at 48kHz, with voice design from a text description, controllable cloning with style guidance, and streaming down to roughly 0.3 RTF on a 4090. Reports 1.84% WER and 75.3% similarity on English Seed-TTS-eval, 0.97% CER on the hard sets and 1.68% average error across 30 languages. Apache-2.0, 36.7k stars.

View repo ↗
3d & spatial

3D & Spatial

3D generation, reconstruction, PBR materials, depth estimation, and spatial computing.

Interactive 3D PortfolioMengTo

Sublevel Studio

A single HTML file with no build step: a Three.js WebGL lobby with a playable arcade cabinet, CRT portfolio screens, custom VHS-style post-processing shaders, procedural textures and camera-state transitions leading into a responsive case-study section, with every image and GLB embedded directly in the page. Hand-authored creative coding rather than anything generated - here as craft, not as an AI release. No licence stated.

View repo ↗
Real-to-Sim ScenesByteDance Seed

Lucida

Composable real-to-sim scene modelling rebuilding an indoor scene as complete, editable object assets rather than one fused mesh, through a parse-generate-place pipeline: parse per-instance evidence from posed RGB, depth, mask and point-cloud observations, generate occlusion-free object images and 3D assets, then place them with GizmoAct, a vision-language model that closed-loop edits object poses inside a 3D editor. Evaluated on R2S-Scene, CA-1M and Aria Digital Twin for detection, pose and Chamfer distance against Boxer, WildDet3D, Any6D, FoundationPose and SAM 3D.

View project ↗
Omni World ModelWorld Labs

Atlas

Announced 1 September into early access: a multimodal autoregressive diffusion transformer using rectified flow, pretrained from scratch to generate, reconstruct and simulate worlds, working natively across text, images, video, 3D data, camera poses and depth maps. Stated capabilities run to a minute of camera-controlled 1440p video, 3D reconstruction from between one and a hundred-plus sparse images, space-time simulation for VFX and robotics, and 360 panorama generation. World Labs reports beating MiniMax H37, Gemini Omni Flash and FLUX 3 on camera control, and Pi3X, VGGT-Omega, Depth Anything 3 and MapAnything on DTU, ETH3D, KITTI and ScanNet.

View release ↗
agents & automation

Agents & Automation

Agent frameworks, orchestration, coding agents, desktop agents, and embodied AI.

Voice Agent Platformdograh-hq

Dograh v1.46.0

Since Week 20 2026 - and that entry recorded nothing because the page would not fetch, so here it is properly. A self-hosted voice-agent platform pitched against Vapi and Retell: visual workflow builder, bring-your-own keys across LLM, STT and TTS or one speech-to-speech model, native MCP support, and telephony through Twilio, Vonage, Telnyx and Plivo. Python, BSD-2-Clause, 5,589 stars, v1.46.0 tagged 3 September.

View repo ↗
Indian Voicebot TelephonyElision Technologies

VoiceLink

Indian telephony built for voicebots rather than adapted to them: mobile, landline, promotional 140-series and transactional 1600-series numbers over SIP and WebSocket, TRAI and VNO licensed, with spam-tag mitigation claimed on the Airtel network. Ships 40+ prebuilt integrations naming Retell AI, Vapi, ElevenLabs and LiveKit. Elision cites 60M+ calls a day across 1,800 enterprise customers; the pricing page is still a placeholder.

View release ↗
OpenClaw 2.0OpenClaw

OpenClaw 2.0

Released 30 August and described by the team as the largest update in the project's history: over 16,000 pull requests from 933 contributors. Installation is rebuilt around existing subscriptions and local models, the browser application is now first-class, cloud sessions are shared for team use, and configuration has moved out of initial setup, alongside changes to messaging, memory, skills, models, automations, security and the native apps. The word "accidentally" in the title is the honest part - the team set out only to fix installation and the browser, and the cleanup cascaded into a full 2.0 over roughly seven weeks against a normal cadence of 106 releases in 230 days.

View release ↗
Multi-Agent Chat Roomssteviebuilds

Agent Room

A localhost-only server on Bun or Node 20 with no external dependencies, creating private rooms where terminal agents such as Codex and Claude Code talk to each other while a human watches or joins from a browser. Long-polling transport, an "only when addressed" mode so agents do not talk over one another, and state kept in ~/.agent-room. MIT, 91 stars. Despite the near-identical name it shares no code with liuyixin-louis/agentroom - different authors, different stacks, different purpose.

View repo ↗
Overnight Agent Runnerkunchenguid

gnhf

Good Night, Have Fun: a TypeScript orchestrator running coding agents overnight, making one small documented git commit per iteration towards a stated objective, with iteration and token caps, rollback and retry on failure, and concurrent agents isolated in git worktrees. Wraps the native CLIs for Claude Code, Copilot, Codex, Cursor, Rovo Dev and OpenCode. MIT, installed with npm i -g gnhf, 3.9k stars.

View repo ↗
Autonomous Web AgencyJackInSightsV2

Automated Agentic AI Web Agency

Fifteen specialised agents on Claude Code running an entire web agency end to end: finding local businesses without websites through Google Places, verifying the lead, writing the copy, building a Vite site, deploying to Vercel, then contacting the owner by email and Bland.ai voice call, following up over Twilio SMS and WhatsApp, closing on a scheduled call and taking payment through Stripe. Bun and Hono over Supabase, with Telegram as the human approval gate and a live agent dashboard. MIT, 41 stars.

View repo ↗
Multi-Agent PlatformLobeHub

LobeHub

Positions itself as a chief agent operator: hiring, scheduling and coordinating agents into groups that run around the clock, over a unified layer across OpenAI, Claude and Gemini models with more than 10,000 MCP-compatible skills and persistent per-agent memory. Next.js, React and TypeScript with Drizzle ORM, under LobeHub's own community licence rather than a standard one. 82.3k stars and 13,509 commits on the canary branch.

View repo ↗
Mobile Automation MCPMobile Next

Mobile MCP

An MCP server giving agents platform-agnostic mobile automation across iOS and Android on simulators, emulators and physical devices. It drives the accessibility tree first and falls back to screenshots, covering device management, app lifecycle, gestures and input, with cloud device access available. TypeScript on Node 20+, Apache-2.0, compatible with Claude Code, Copilot, Gemini, Cline and Cursor. 6.3k stars.

View repo ↗
Agent Control CentreOpenHands

OpenHands v1.16.0

A self-hosted control centre for coding agents - its own, plus Claude Code, Codex, Gemini or anything ACP-compatible - running against local, Docker, VM or cloud backends with any model behind them. The new Automation Server adds scheduled and webhook-triggered workflows wired into Slack, GitHub and Linear. TypeScript, MIT, 86.3k stars, still labelled beta at v1.16.0.

View repo ↗
Open Drone Platformswarm-subnet

Langostino

An open autonomous drone reference platform pairing custom INAV flight-control firmware with a ROS2 Humble autonomy layer on a Raspberry Pi, over LiDAR, GPS and IMU, with the decision nodes written in Python. The point is that it ships the assembly guide, the bill of materials and the setup docs, so aerial autonomy becomes something you can build rather than a black box you buy. MIT, versioned 1.0.0, 366 commits.

View repo ↗
LinkedIn Agent Skillssergebulaev

LinkedIn Skills

Eleven Claude Code and Codex skills for LinkedIn work: post writing with more than 20 hook formulas, comment and reply drafting, auditing a post against algorithm and AI-detection patterns, an AI-tell humaniser, hook extraction, a seven-day planner, engagement monitoring, profile optimisation, employee advocacy planning and cross-platform repurposing. Python, with optional Apify for reading data and Publora for publishing, and nothing publishes without explicit human approval. MIT.

View repo ↗
developer tools

Developer Tools

SDKs, inference engines, data pipelines, monitoring, and AI infrastructure.

Convex Social Componentzernio-dev

convex-zernio

A Convex component in TypeScript letting a Convex app schedule and publish social posts through the Zernio API. Submission is durable through a workpool with retries and rate limiting, per-platform status comes back over HMAC-verified webhooks, and a React hook subscribes to it. Profiles are multi-tenant and test mode is the default, so posts land as drafts rather than going live. Apache-2.0, last pushed 30 August, still small at 2 stars.

View repo ↗
Idea Reality Checkermnemox-ai

idea-reality-mcp

An MCP server giving a coding agent a reality check on a startup idea by scanning GitHub, Hacker News, npm and PyPI, returning a 0-100 score with competitor evidence and trend direction. Quick mode covers GitHub and HN in under three seconds, deep mode all four. Product Hunt was dropped because its API has no text search. Python, MIT, 813 stars, installed via uvx or Smithery - though the repo has been in maintenance mode since August.

View repo ↗
Unified Social APIPostproxy

Postproxy

One publishing and engagement API across Instagram, TikTok, LinkedIn, X, YouTube, Facebook, Threads, Bluesky, Pinterest, Telegram and Google Business, where a single payload handles per-platform formatting alongside comment, DM and review management, webhooks and status logging. Node, Python and Ruby SDKs, a sandbox, verified n8n, Make and Zapier integrations, and an MCP server for agent workflows. States 35ms average response, EU-hosted, SOC 2 Type II in progress; free tier, no card.

View release ↗
Agent Skill ScannerNVIDIA

SkillSpector v2.11.0

Since Week 24 2026: a security scanner for agent skills across Claude Code, Codex CLI, Gemini CLI and MCP, matching 71 vulnerability patterns in 17 categories - prompt injection, data exfiltration, privilege escalation, supply-chain risk - through fast static scanning, an optional LLM semantic pass, live OSV.dev CVE lookups and a 0-100 risk score. It was built on an empirical study of 42,447 skills that found 26.1% carrying vulnerabilities and 5.2% showing likely malicious intent. Apache-2.0, 16.3k stars, v2.11.0 tagged 28 August.

View repo ↗
Codebase Diagram Generatortt-a1i

Archify v2.16.0

An agent skill for Cursor, Claude Code, Codex CLI and OpenCode turning a codebase or a written system description into an interactive system map - architecture, workflow, sequence, data-flow and lifecycle diagrams - rendered as self-contained HTML and SVG through a typed JSON intermediate representation. Includes an Architecture Delta view diffing before and after, route tracing, and export to PNG, SVG or WebM. MIT, 49.6k stars, v2.16.0 tagged 30 August.

View repo ↗
Peer-to-Peer InferenceSwarmLLM

SwarmLLM

Browser-based peer-to-peer inference splitting one model across every device in a room so they act as a single compute unit, running a model no single device could hold. Devices join on a four-letter room code with no install, account or server; each downloads a slice sized to its capability and the response is computed collectively and appears on every screen at once. The site names Qwen 3.8 as the model and argues the privacy case that nothing leaves the room. No organisation, version or pricing appears anywhere on the page.

View release ↗
GPU Architecture GuideDr Ashish Bamania

NVIDIA GPUs for AI Engineers

A free explainer on the GPU hardware an AI engineer actually has to reason about: streaming multiprocessors, CUDA and tensor cores, on-chip SRAM against HBM, and the NVIDIA line from Volta and the V100 in 2017 through T4, A100, H100, H200, L40S, B100, B200, B300 and Rubin to Feynman in 2028. It then covers interconnects - PCIe, NVLink 4 to 6, NVSwitch, NVLink-C2C - and scales outward through Spectrum-X Ethernet and Quantum InfiniBand to the node, rack, cluster and facility hierarchy, mapping tensor, data and pipeline parallelism onto it.

View release ↗
GPU Concepts PrimerDr Ashish Bamania

GPU Concepts for AI Engineers

Nine GPU concepts in four groups: computation hardware covering streaming multiprocessors, CUDA cores across INT32, FP32 and FP64 and tensor cores doing matrix multiply-accumulate; GPU memory with HBM holding weights and activations above an on-chip hierarchy; program execution; and multi-GPU connectivity. The framing is debugging - working out where the compute budget actually goes when a model runs slower than expected. The first four concepts are free to read, five through nine sit behind the paywall.

View release ↗
LLM Inference InternalsDr Ashish Bamania

A Hardware-Level Tour of LLM Inference

Traces inference down to the metal: weights moving from SSD through CPU DRAM into GPU HBM over PCIe, the CPU-HBM-SRAM hierarchy, and the tensor, pipeline, context, expert and data parallelism used to spread work across GPUs. It separates the compute-bound prefill phase, where the whole prompt is processed at once, from the memory-bound decode phase producing one token at a time, and explains the KV cache as what prevents the recomputation. The argument it builds to is that memory bandwidth, not raw compute, is the wall.

View release ↗
Distributed Data ParallelDr Ashish Bamania

Distributed Data Parallel Explained

A free, diagram-led walk through PyTorch DDP in eleven steps: replicating the model per GPU, splitting the dataset with DistributedSampler, the forward and backward passes, and gradient synchronisation through AllReduce, plus the optimisation overlapping gradient computation with communication. A worked three-GPU example carries it. It closes on the limit that motivates model parallelism - an 8B Llama needs roughly 151GB, more than an 80GB H100 holds. Billed as Part 1, with the code in Part 2.

View release ↗
PyTorch DDP WalkthroughDr Ashish Bamania

Train a CNN with PyTorch DDP

The hands-on Part 2 to the DDP explainer: training a CNN image classifier across several NVIDIA T4 GPUs on Kaggle with PyTorch DistributedDataParallel, over CIFAR-10's 60,000 32x32 images in ten classes. It walks the imports for torch.distributed and TorchVision, then a three-block convolutional architecture with batch normalisation, pooling and fully connected classification, plus the scaffolding for the distributed helper functions. Marked paid, so only the preview is readable without a subscription.

View release ↗
Llama Distributed TrainingDr Ashish Bamania

Distributed Training of Llama Explained

How Meta trains at Llama scale using 4D parallelism: FSDP sharding parameters, optimiser states and gradients rather than replicating them; context parallelism splitting 128K+ token sequences; tensor parallelism splitting weights along the hidden dimension; and pipeline parallelism splitting layers into stages fed by micro-batches, with expert parallelism as a fifth dimension for MoE. Grounded in Llama 3-405B's real run - 16,000 GPUs at 80GB HBM3, 54 days, 30.84M GPU-hours - and it gives the arrangement priority order as TP, CP, PP then DP and FSDP, ranked by communication bandwidth.

View release ↗
Agent Visualiserliuyixin-louis

AgentRoom

A Tauri desktop app - Rust backend, React and Canvas 2D front end - drawing your running coding agents as pixel-art characters in a virtual office, with a work room for active agents and a break room for idle ones. It tails the JSONL transcripts Claude Code, Codex and Gemini write and turns them into events driving character state and pathfinding, alongside full-text session search, per-project offices, a token-usage dashboard, AI session tagging and one-click iTerm2 integration. MIT. Not the same project as the similarly named agent-room, despite the collision.

View repo ↗
Free Market Data ArchiveLondon Strategic Edge

London Strategic Edge

Not the consultancy the domain suggests: a free financial-data platform hosting what it calls the largest free market-data archive, 133 billion ticks across 118,000 datasets covering stocks, options, forex, crypto, futures and macro series back to 1900 for over 100 countries. It ships working tools on top - strategy backtesting, LSTM and XGBoost forecasting, live charts, options analysis and screening across 16,000+ instruments - with history downloadable as Parquet or CSV through a free API key, no paywall and no card.

View release ↗
Engineering Workflow PluginLauren Tan

pstack

A Cursor plugin, with an unofficial Claude Code port, expanding a short request into a structured engineering workflow - 23 workflow skills, 21 engineering principles, 22 task playbooks and specialised subagents - entered through a single /poteto-mode command. The skills use the standard SKILL.md format so they carry across Claude Code, Codex and other agents, and sub-tasks can be routed to different models. Distributed through Cursor's official plugin system.

View release ↗
Self-Hosted Financesecuro-finance

Securo

Self-hosted personal finance manager on FastAPI, SQLAlchemy, Alembic and Celery with a React and Vite front end over PostgreSQL and Redis. Handles multi-account and multi-currency tracking, file import, auto-categorisation rules, budgets and savings goals, bank sync through Pluggy, Enable Banking and SimpleFIN, and login by 2FA, OIDC or passkey, with optional self-hosted AI agents that can call tools. AGPL-3.0, 3.1k stars.

View repo ↗
Debugger MCP Serverwasdubya

x64dbgMCP

An MCP server in C++ and Python wiring an LLM into the x64dbg debugger so reverse engineering can be driven in natural language, exposing over 40 SDK functions - register queries, byte-pattern search and the rest - across 32-bit and 64-bit targets. GPL-3.0, 517 stars. The debugger underneath is a decade-old project with nothing new this cycle; the MCP layer is what is new.

View repo ↗
AI Slide DecksStackBlitz

Bolt Slides

Lets an agent generate an interactive presentation from one prompt, where every slide is a live React component rather than an image - so a slide can hold real data, 3D elements or a working prototype. Ships presenter mode with synced tabs, annotations and deep linking, more than 20 prebuilt slide components covering charts, quotes and pricing layouts, CSS-token theming, and a bundled skill guide telling an agent how to compose them. TypeScript and Vite, MIT, 924 stars.

View repo ↗
Local Model Servermagnitudedev

Magnitude

A local inference server that profiles the machine it is on, recommends and downloads models that will actually fit, then plugs them into existing coding agents including Pi, OpenCode, Hermes and Cline so they run free, private and offline. Models load just in time and unload when idle. TypeScript and Bun, Apache-2.0, on macOS, Linux and Windows through WSL, installed with npm i -g @magnitudedev/cli. 3.3k stars.

View repo ↗
Order Routing Platformopenshiporg

Openship

Order routing between sales channels - Shopify, WooCommerce, eBay, Amazon - and fulfilment partners including supplier stores, 3PLs, email and Google Sheets, modelled as shops, channels, links and product-level matches so orders route automatically. Next.js 15 and KeystoneJS 6 behind a GraphQL API over PostgreSQL and Prisma. AGPL-3.0, 1.6k stars. E-commerce infrastructure rather than an AI release.

View repo ↗
Stickman Prompt Skillkaomei

Stickman Video Director

A Codex skill rather than a model: it turns an article or piece of copy into a complete one-minute stickman animation production package for Gemini Omni Flash - confirmed narration, a directorial proposal, and six independent prompts carrying timing beats, camera moves, transitions and sound design. Light and dark themes, 9:16, 16:9 and 1:1. MIT, 769 stars. It writes the prompts; another model renders the video.

View repo ↗
Multivariate ForecastingGoogle Research

TimesFM-3

A 330M decoder-only transformer for time-series forecasting pretrained on over a trillion time points, and the first in the line that is natively multivariate - where TimesFM and 2.5 were strictly univariate, this handles multiple targets, past covariates and past-future dynamic covariates in one forward pass. Uses 32-step patches, alternating causal temporal and full variate attention, and non-autoregressive contiguous patch masking to emit nine quantiles per target in a single pass. Tops Gift-Eval, FEV-Bench and TIME against Chronos-2, Toto 2.0 and TimesFM-2.5 on both point and probabilistic metrics; code and weights released, BigQuery integration promised.

View release ↗
Video Delta AttentionUC Berkeley, Impossible and UT Austin

VDN

Video DeltaNet is a hybrid attention design pairing sliding-window softmax attention with a linear-attention branch, aimed at the more than 85% of runtime softmax attention eats in video diffusion, without the identity, layout and temporal-consistency damage linear attention usually brings. Applied to MiniMax H3 the authors report 2.65x per layer - 125.3ms against 332.5ms on one B200 - a 14.3-second 768p clip in 11.23 seconds on eight B200s, and up to 74.5x against a dense single-GPU baseline. Code and weights released.

View project ↗
VDN ComfyUI NodeSaganaki22

ComfyUI-VDN-H3

Ports VDN into ComfyUI for MiniMax-H3, patching attention at runtime - exact softmax for nearby frames, a constant-cost recurrent linear branch for distant temporal context - without touching ComfyUI core. Needs the H3 base weights plus a Qwen3VL-32B text encoder and separate video and audio VAEs. On a single RTX 5090 at int8, 1280x736 and 145 frames it runs about 17 seconds an iteration, roughly two and a quarter minutes total; the paper's headline 74.5x needs eight GPUs and kernels consumer hardware does not have. The node is Apache-2.0, the weights gated under the MiniMax H3 community licence.

View repo ↗
research & safety

Research & Safety

Research papers, alignment, red-teaming, interpretability, and benchmarks.

Skill EvolutionTang, Rashtchian, Ferng, Tomkins, Juan and Vu

WikiSkill

Compiling Agent Experience into Persistent Knowledge for Skill Evolution proposes co-evolving an agent's skills alongside a wiki-style knowledge base, keeping raw execution experience, accumulated knowledge and executable skills separate so consolidated experience feeds the next skill update rather than being replayed each time. The paper reports consistently beating state-of-the-art skill-evolution methods across benchmarks and models. The ablations carry the interest: larger models gain more from evolved skills, a smaller model with skills can beat a larger one without them, and skills transfer across model families.

Read paper ↗
AI Adoption LadderEvery

The Eight Levels of AI Adoption

A maturity ladder running Chatbot, Copilot, Agent with human approval checkpoints, Autopilot, Workflows, Assistant doing proactive unprompted background work, Multi-agent managing several long-running agents at once, and Orchestrator where manager agents coordinate teams of sub-agents. Its useful claim is that higher is not better: the right rung depends on how much you trust the model's accuracy and what an error costs. It places most knowledge workers at levels one to four and engineers at five to eight.

View release ↗
Efficient ARC-AGI ModelMithil Vakde

mdlARC

A 75M-parameter transformer trained from scratch for about 67 cents reaches 44% on ARC-AGI-1 - no LLM involved, and among the best low-cost scores on that benchmark. Ships the full dataset-building pipeline and three training intensity modes, with a run taking roughly two hours on an RTX 5090 over PyTorch, numba and flash-attn. MIT. Set against everything else that shipped this week, the most quietly striking result of the seven days.

View repo ↗
WeatherNext 3Google DeepMind

WeatherNext 3

A functional generative network mesh transformer ingesting real-time geostationary satellite mosaics alongside historical analysis, forecasting surface variables at 5km - five times sharper than WeatherNext 2's 25km - and refreshing hourly instead of every six hours, with new 100m turbine-height wind and clean-energy variables. Google reports CRPS precipitation gains up to 60% against IMERG, 30% against MRMS and 10% against rain gauges at early lead times, and cites independent live evaluation by Brightband. Going into Search, Gemini, Maps, the Maps Platform Weather API and Earth Engine, with data in BigQuery.

View release ↗
Fly Brain ConnectomeGoogle Research and HHMI Janelia

Male Fruit Fly Connectome

Published in Cell on 3 September: the complete central nervous system of the male fruit fly - central brain, optic lobes and ventral nerve cord - at over 166,000 neurons and 125 million synaptic connections, the largest brain map by neuron count so far. Built from electron-microscopy sectioning, AI 3D reconstruction using flood-filling networks, and a PATHFINDER system that puts synthetic neurons into the training data, then verified by hand with Columbia and Harvard. With the female map already done it makes direct male-female comparison possible, and is positioned as a step towards vertebrate brains.

View release ↗