AI models and tools worth knowing about
Everything the AI channels we follow have named, each with a plain description of what it is, how you get it and what it needs to run.
128 items from 11 videos · ✓ marks the ones we have tried ourselves. ·
RSS feed · rebuilt 2026-09-04
2026-09-03 AI Search
Claude Fable 5.1 is savage
Claude Fable and Mythos 5.1Anthropic's newest pair of general-purpose assistant models, aimed at coding and knowledge work.
paid service · needs none, runs in the cloud
2026-08-30 AI Search
Ox Alpha reveal, realtime Minimax, Qwen Next, Hy4, robot olympics: AI NEWS
VoiceMemBuilds a memory of a spoken conversation — who was mentioned, what was said, and the speaker's preferences — from the audio itself.
open source
S1A robot control model that learns a new task from a single video demonstration, with no retraining.
hosted service
Qwen 3.8 Flash NextAlibaba's fast, low-cost tier of its Qwen assistant, announced alongside its chat and image tools.
hosted service
OrbitA benchmark for testing how well camera-position reconstruction works on real 360-degree video.
research project
One Video One WorldRebuilds a moving 3D scene from one ordinary video, with separate objects you can simulate, and no training step.
research project
Omni 1.1 FlashGoogle's video generation model for developers, with added controls over how a shot is composed.
paid service · needs none, runs in the cloud
Minimax FastH3A faster rebuild of MiniMax H3 that generates video and its audio together.
open source
Hy4Language model (hy_v4, text-generation) published on Hugging Face as tencent/Hy4-preview.
open, apache-2.0
H3 MaxA speed-tuned version of MiniMax H3, run on fal's own servers rather than your machine.
paid service · needs none, runs in the cloud
Google PPEGoogle's system for automatically building global environmental prediction models from Earth observation data.
research project
GLM 5.3 FlashThe announcement page for the faster, cheaper tier of Zhipu's GLM 5.3. The page returned no readable text to us.
unclear
GLM 5.3Language model (glm_moe_dsa, text-generation) published on Hugging Face as zai-org/GLM-5.3.
open weights under the publisher's own licence — read it
Gemini 3.5 TranscribeGoogle's speech-to-text service, updated to handle messier audio and context.
paid service · needs none, runs in the cloud
FixAnythingCleans up rendering errors in 3D scenes by passing them through a video model.
research project
Fibo 1.5Image generator (art, text-to-image) published on Hugging Face as briaai/Fibo-1.5.
open weights, access must be requested
DiffusionOPSDA training technique for image models that makes their fine-tuning cheaper to run and easier to inspect.
research project
Code World ModelA coding agent that keeps an internal model of the system it is working on.
research project
Block 3DGenerates 3D objects from a text prompt, building them block by block to save compute.
research project
2026-08-25 Mickmumpitz
We Open Sourced World Generation
Wan 2.1Wan: Open and Advanced Large-Scale Video Generative Models.
open source, apache-2.0
UniSHARP Project PageInsta360's research on turning a single photograph into a sharp, viewable 3D scene.
research project
UniSHARP (Insta360)Official implementation of UniSHARP: Universal Sharp Monocular View Synthesis.
source published, licence unclear
TritonDevelopment repository for the Triton language and compiler.
open source, mit
SeedVR2Repo for SeedVR2 (ICLR2026) & SeedVR (CVPR2025 Highlight).
open source, apache-2.0
SageAttention[ICLR2025, ICML2025, NeurIPS2025 Spotlight] Quantized Attention achieves speedup of 2-5x compared to FlashAttention, without losing end-to-end metrics across language, image, and video models.
open source, apache-2.0
Postshot (Jawset)Postshot, a Windows desktop application for building photorealistic 3D reconstructions from footage.
unclear, the site carries a pricing page · needs Windows 10 or later, Nvidia RTX 2060 or better
NVIDIA LyraProject Lyra: Open Generative 3D World Models.
open source, apache-2.0
MoGe (Microsoft)[CVPR'25 Oral] MoGe: Unlocking Accurate Monocular Geometry Estimation for Open-Domain Images with Optimal Training Supervision.
source published, licence unclear
Matrix-3D Project PageTurns one 360-degree photograph into a 3D scene you can move around in. Project page.
open source
Matrix-3D PaperThe research paper behind Matrix-3D, on generating explorable 3D worlds from a single panorama.
paper only
✓Matrix-3D (Skywork AI)Generate large-scale explorable 3D scenes with high-quality panorama videos from a single image or text prompt.
open source, mit
Marble (World Labs)A commercial tool for creating and sharing 3D worlds in the browser.
hosted service · needs none, runs in the cloud
LightX2VLightweight Image Video Action Generation Inference Framework.
open source, apache-2.0
LichtFeld StudioTrain, inspect, edit, automate, and export 3D Gaussian Splatting scenes from a single native application.
open source, gpl-3.0
Krea 2 Technical ReportKrea's image generation models, released with downloadable weights in a standard and a fast version.
open weights
HY-World 2.0 (Tencent Hunyuan)HY-World 2.0: A Multi-Modal World Model for Reconstructing, Generating, and Simulating 3D Worlds.
source published, licence unclear
Brush3D Reconstruction for all.
open source, apache-2.0
Apple SHARP Project PageApple's method for turning one photograph into a viewable 3D scene in under a second, with paper and code published.
open source
Apple SHARP PaperApple's paper on reconstructing a viewable 3D scene from a single photograph in under a second.
paper only
Apple SHARPSharp Monocular View Synthesis in Less Than a Second.
source published, licence unclear
2026-08-23 AI Search
New AI waifus, new Deepseek, realtime worlds, Happy Shrimp, tiny TTS: AI NEWS
SenseNova U1.5An 8-billion-parameter reasoning model from SenseTime, distributed through ModelScope.
open weights
Qwen Video EditEdits video by written instruction, including long clips where each segment gets a different instruction.
research project
Ornith 1.5A model family built to improve its own answers in a loop rather than in a single pass.
open source
Happy ShrimpA music service that writes a complete song, lyrics and vocals included, from a one-line prompt.
hosted service · needs none, runs in the cloud
GeoWeaverBuilds a consistent 3D reconstruction from long video sequences without drifting off course.
research project
Gen 1.5A robot model that picks up a new task from a few seconds of demonstration, or a few minutes of data.
hosted service
EvokeNamed in the video, but the project page it linked to no longer exists.
unavailable
Deepseek visionDeepSeek's image understanding: send a picture with your question and it reads charts, screenshots and text.
paid service · needs none, runs in the cloud
Comfy MCPLets a coding assistant drive a ComfyUI install on your own machine, using your GPU, your models and your custom nodes.
open source · needs your own GPU
Bernini v2Model weights (bernini, image-text-to-video) published on Hugging Face as ByteDance/Bernini-Diffusers-v2.
open, apache-2.0
✓AvoNVIDIA's argument that the scaffolding around a model, not the model alone, decides how well an agent performs on long tasks.
paper only
Audio8 TTS 0.1BSpeech synthesiser (arktts, feature-extraction) published on Hugging Face as Audio8/Audio8-TTS-Preview-0.1b.
open weights under the publisher's own licence — read it
4DAnyoneRebuilds a moving 3D person from one ordinary handheld video, with no camera calibration.
research project
2026-08-16 AI Search
New Deepseek, GLM 5.3, Grok 4.6, LTX 2.5, Qwen 3.8, Gemini 3.7: AI NEWS
WorldClawTurns a single open-ended prompt into a 3D world you can walk through and edit.
research project
Sign to textDeepMind's sign-language-to-text model, now driving features in shipping products.
hosted service
ScopeGives an existing video model precise camera control by telling it which direction each part of the frame is being viewed from.
research project
Qwen 3.8 maxLanguage model (qwen3_5_moe_text, text-generation) published on Hugging Face as Qwen/Qwen3.8-2.4T-A95B.
open weights under the publisher's own licence — read it
✓Qwen 3.8 27BVision-language model (qwen3_5, image-text-to-text) published on Hugging Face as Qwen/Qwen3.8-27B.
open, apache-2.0
Nemotron LightningA small open NVIDIA model plus a routing library, meant to run the same workload across a laptop, a workstation or a data centre.
open weights
Muse GlimmerMeta's 30-billion-parameter model for agents that run continuously on your own machine, tuned for tool use and recovering from mistakes.
open weights, apache-2.0 · needs a single GPU
MiDashengGenerates full audio scenes of variable length by treating sound the way a language model treats words.
research project
MatrAIxTesting infrastructure that simulates a range of users against your product before you release it to real ones.
open source
Magi 2Model weights (text-to-video, image-to-video) published on Hugging Face as sand-ai/MAGI-2-preview.
open, apache-2.0
LTX 2.5Generates multi-shot video in one pass, edits real footage, and exports at cinema quality. Weights are downloadable.
open weights · needs runs on your own hardware
JoyAI video edit[Official Repo] JoyAI-Video-Edit: Real-Time Open-Ended Video Editing with Autoregressive Diffusion.
open source, apache-2.0
Index TTS 2.5Speech synthesiser (indextts, text-to-speech) published on Hugging Face as IndexTeam/IndexTTS-2.5.
open weights under the publisher's own licence — read it
Grok 4.6xAI's assistant model, this version aimed at agents that run for a long time without supervision.
paid service · needs none, runs in the cloud
GPT UltrafastAn OpenAI service tier that runs their model far faster on specialised chips, quoted at up to 750 words per second.
paid service · needs none, runs in the cloud
GLM 5.3The announcement page for Zhipu's GLM 5.3 language model. The page returned no readable text to us; the downloadable weights are listed separately in this archive.
unclear
Gemini 3.7 FlashGoogle's mid-tier Gemini model, positioned for everyday coding and agent work.
paid service · needs none, runs in the cloud
Dyna 2A robot control model trained on over a million hours of ordinary human video, presented with evidence that it keeps improving with scale.
hosted service
Deepseek V4 0813Language model (deepseek_v4, text-generation) published on Hugging Face as deepseek-ai/DeepSeek-V4-Pro-0813.
open, mit
Cactus NeedleA tiny 45-million-parameter model for calling tools and pulling structured data out of text, small enough to ship inside an app.
open weights · needs runs in about 28 MB of memory
2026-08-15 AI Search
The BEST local AI music generator is here!
MuscriptorThe link points at a publisher's Hugging Face account rather than a specific release, and the page could not be read without signing in.
unclear
Minimax Music for ComfyOfficial instructions for running MiniMax Music 3 inside ComfyUI to produce complete songs up to five minutes long.
free to use
2026-08-12 AI Search
The BEST local AI video generator just got BETTER!
Smaller modelsModel weights published on Hugging Face as Kijai/MiniMax-H3-experimental.
unclear
Realism loraModel weights (minimax-h3, lora) published on Hugging Face as fal/MiniMax-H3-Realism-People-LoRA.
open weights under the publisher's own licence — read it
MiniMax H3 Turbo LoraVideo generator (minimax-h3, text-to-video) published on Hugging Face as larryvrh/MiniMax-H3-Turbo-Lora.
open, apache-2.0
MiniMax H3 TurboVideo generator driven by a still image (minimax-h3, comfyui) published on Hugging Face as joyfox/MiniMax-H3-Turbo.
unclear
Minimax h3 TurboVideo generator driven by a still image (t2v, i2v) published on Hugging Face as lightx2v/Minimax-h3-Turbo.
open, apache-2.0
MiniMax H3 prompt guide (reference)Model weights (minimax-h3, text-to-video) published on Hugging Face as MiniMaxAI/MiniMax-H3.
open weights under the publisher's own licence — read it
MiniMax H3 prompt guide (basic)Model weights (minimax-h3, text-to-video) published on Hugging Face as MiniMaxAI/MiniMax-H3.
open weights under the publisher's own licence — read it
MiniMax H3 GGUFVideo generator driven by a still image (gguf, comfyui) published on Hugging Face as Abiray/MiniMax-H3-GGUF.
open weights under the publisher's own licence — read it
MiniMax H3 comfyModel weights published on Hugging Face as Kijai/MiniMax-H3_comfy.
unclear
MiniMax H3Model weights (minimax-h3, text-to-video) published on Hugging Face as MiniMaxAI/MiniMax-H3.
open weights under the publisher's own licence — read it
Live previewModel weights published on Hugging Face as Kijai/MiniMax-H3-TAE.
open, apache-2.0
Image workflowModel weights published on Hugging Face as reverentelusarca/minimax-h3-comfyui-workflows.
open weights under the publisher's own licence — read it
GGUFsModel weights (gguf, text-to-video) published on Hugging Face as unsloth/MiniMax-H3-GGUF.
open weights under the publisher's own licence — read it
Easy workflowThe easiest way to use MiniMax H3. One compact workflow for T2V, I2V, first/last-frame, and reference video generation, with a unified multi-media input, powerful @ references, and inline dialogue blocks. Less wiring. More creative control.
open source, mit
2026-08-05 AI Search
The BEST local AI video generator is here! Minimax H3 tutorial
Workflows and filesOfficial ComfyUI instructions for MiniMax H3, covering text, image and reference-driven video with stereo sound.
free to use
SpectrumTraining-free Spectrum acceleration for ComfyUI’s native MiniMax H3 audio-video model. Uses Chebyshev ridge feature forecasting to skip selected H3 transformer evaluations, with adaptive scheduling, sampler-aware support for Euler, ER-SDE, RES, SEEDS and SA-Solver, CPU/VRAM history storage, and fail-closed native fallbacks.
open source, gpl-3.0
Sage attention wheelsFork of SageAttention for Windows wheels and easy installation.
open source, apache-2.0
More on licenseModel weights (minimax-h3, text-to-video) published on Hugging Face as MiniMaxAI/MiniMax-H3.
open weights under the publisher's own licence — read it
✓Minimax H3 modelsModel weights (diffusion-single-file, comfyui) published on Hugging Face as Comfy-Org/MiniMax-H3.
open weights under the publisher's own licence — read it
KJ nodesVarious custom nodes for ComfyUI.
open source, gpl-3.0
2026-08-02 AI Search
New Deepseek, Seedance 2.5, Minimax H3, Gemini Robotics, AMD models: AI NEWS
WonderA world model that generates an interactive scene from a starting image or video.
research project
✓Seedance 2.5ByteDance's video model that generates picture and sound together, built for 30-second stories with reference images and editing.
paid service · needs none, runs in the cloud
RedesignOfficial repository for "ReDesign" (ECCV 2026).
source published, licence unclear
PrismA compact way of representing which of a robot's own signals actually matter for the task it is doing.
research project
Phi ZeroA world model that learns how physical scenes change by watching ordinary video.
research project
✓Minimax H3An all-in-one generation model that reads text, images, video and audio together, and produces video with stereo sound at up to 2K and 15 seconds.
open weights
Kimi K3Vision-language model (kimi_k3, feature-extraction) published on Hugging Face as moonshotai/Kimi-K3.
open weights under the publisher's own licence — read it
✓InstellaAMD's fully open 16-billion-parameter language model, of which only 2.8 billion are used per word, trained on AMD's own accelerators.
open weights
Inkling SmallVision-language model (inkling_mm_model, image-text-to-text) published on Hugging Face as thinkingmachines/Inkling-Small.
open, apache-2.0
Ideogram obj removerAn online tool for removing unwanted objects from a photograph. The page blocks automated readers, so nothing beyond the name was confirmed.
hosted service · needs none, runs in the cloud
✓ID V2VRestyles a video into a different look while keeping the person recognisably themselves. From Netflix and Eyeline Labs.
research project
Gemini voice typingDictation in the Gemini app for macOS: speak, and it produces cleaned-up text, edits and summaries.
hosted service · needs none, runs in the cloud
Gemini Robotics 2DeepMind's robot model extending control from the hands to the whole body.
research project
✓Deepseek V4 Flash 0731Language model (deepseek_v4, text-generation) published on Hugging Face as deepseek-ai/DeepSeek-V4-Flash-0731.
open, mit
✓Crisper WhisperSpeech recognition that transcribes exactly what was said, in many languages, with the timing of each individual word.
open weights, non-commercial only
2026-07-26 AI Search
Claude Opus 5, GPT 6 hack, Flux 3, new Gemini, quantum breakthrough, new Qwen: AI NEWS
✓ShotPlanGenerates multi-shot video with clean cuts and camera moves at the exact frames you ask for.
open source
Sana Video 2NVIDIA's efficient 720p video generator, built to cut the cost of the attention step.
research project
Qwen 3.8The paid subscription page for Alibaba's Qwen cloud service.
paid service · needs none, runs in the cloud
Opus 5Anthropic's top assistant tier, aimed at agents that run for long stretches and at professional coding work.
paid service · needs none, runs in the cloud
OpenDreamerAn openly released world model with a playable demo, published together with an account of how it was trained.
open source
OpenAI hackOpenAI and Hugging Face's joint account of a security incident during model evaluation. News, not a tool.
news article
New Gemini modelsThree new Gemini tiers, including a lighter one and a security-focused one.
paid service · needs none, runs in the cloud
✓NanbeigeA 3-billion-parameter reasoning model, small enough for a single consumer graphics card.
open weights · needs one consumer GPU
Mage FlowNamed in the video, but the Hugging Face page could not be read without signing in, so it may be restricted or withdrawn.
unclear
Laguna S2.1poolside's coding model, this version aimed at work that takes many steps to finish.
hosted service · needs none, runs in the cloud
HomiePuts a specific person and object into generated video while keeping both recognisable.
research project
GPT live in desktopVoice conversation in the ChatGPT desktop app.
hosted service · needs none, runs in the cloud
Google quantum breakthroughGoogle's work on using reinforcement learning to improve quantum error correction. News, not a tool.
news article
GLM with visionVision-language model (sglang, glm5v) published on Hugging Face as baseten/GLM-5.2-Vision-NVFP4.
open, mit
Flux 3Black Forest Labs' single model covering images, video, audio and predicting actions.
unclear
ChatGPT HealthLets eligible US users connect medical records and Apple Health to ChatGPT for personalised health summaries.
hosted service · needs none, runs in the cloud