Multimodal
Systems that combine text, image, audio, video, or UI inputs.
Stories
Filter storiesOpenAI introduced GPT-Image 2.5 Flare and Sunburst API models alongside ChatGPT Images 2.5. The release emphasizes lower generation latency, stronger multi-turn edits, and better reference-subject preservation.
World Labs says Atlas combines visual generation with scene reconstruction. It can reconstruct scenes from images, generate camera-controlled frames, and reframe video.
Meta launched Muse Voice Transcribe, combining streaming transcription, speaker diarization, and endpointing in one model. Meta says it supports code-switching and costs $0.18 per hour.
Google says Gemini Omni 1.1 Flash can reference up to 10 seconds of preceding video when extending a scene, rather than only the final second. It also adds keyframe interpolation and object editing for more controlled video generation.
Perceptron released weights, inference code, and training details for Isaac 0.5, an embodied model for video perception, reasoning, and robot control. The 36B dynamic-MoE model was trained on 1 million hours of video.
WAN 3.0 is now available through Replicate, OpenRouter, and ComfyUI for text-, image-, and reference-driven video generation up to 1080p. Pika says the model supports up to 20 references and 30-second output, while quality and price comparisons remain vendor-reported.
Meta published Muse Spark 1.2 evaluations covering tool-based web page and game creation, robotics planning, and audio-visual tasks. A separate result places it first on Design Arena’s video-to-website benchmark.
OpenRouter is testing Ox Alpha, a free stealth model with a 1M-token context window. The model accepts text, image, and video inputs and is offered with zero data retention during the test.
Sentence Transformers 6.0 adds MultiVectorEncoder for ColBERT-style training, inference, and interpretation, including visual-document retrieval. Index size can increase substantially: 4,874 passages expanded to 608,414 token vectors in one example.
Qwen released its open-weight Qwen3.8 27B vision-language model with 262K native context and adjustable reasoning. In a 484-sample test, enabled-thinking scores fell from above 92% through 64K to 74.3–81.8% at 128K.
LocalLLaMA users said Qwen 3.8 Max 2.4T open weights are text-only while the API keeps vision support. A linked Qwen3.8-27B ModelScope page reportedly returned 404 before release.
Google opened Gemini Omni Flash to developers for video creation and editing from text, image, video, or audio references. The rollout spans Gemini, Flow, AI Studio, the Gemini API, and Vertex AI, with supporting posts showing editing demos rather than independent performance tests.
ComfyUI says Seedance 2.5 is live via Partner Nodes with 30-second runs, up to 50 references, timeline shot control, editing, and multilingual lip sync. Fal, Pika, Venice, and other tools also added access.
MiniMax H3 now has Hugging Face weights, fal endpoints, AI Toolkit LoRA support, and reported single-RTX-5090 local runs. MiniMax also said deployment in the US, EU, UK, and South Korea is available through formal authorization.
MiniMax released H3 weights on Hugging Face for text-to-video, image-to-video, reference-to-video, and editing workflows. vLLM-Omni, ComfyUI, SGLang Diffusion, and fal added support at launch.
Google released Gemini Robotics 2, ER 2, and On-Device 2 for humanoid control, embodied reasoning, and on-device adaptation. Demos showed sub-second streaming and multi-robot task handoffs.
Thinking Machines released Inkling-Small, a 276B-parameter MoE with 12B active parameters, multimodal inputs, and a 1M-token context window. Providers added day-zero vLLM, SGLang, Modal, and gateway support.
MiniMax announced H3 for app and API access, with OpenRouter, Vercel, Pika, fal, Runway, and Venice integrations. Artificial Analysis ranked it No. 1 for video editing, and MiniMax said open weights are planned.
Moonshot published Kimi K3 weights, a technical report, and a blog for a 2.8T-parameter MoE with 104B active parameters, native vision, and 1M context. The license adds separate terms for large model-as-a-service providers.
Alibaba opened Qwen 3.8 Max Preview on Alibaba Cloud, Qwen Chat, Qoder and the web, describing it as a 2.4T model headed for open weights. Early testers praised vision results but disputed coding claims.
Thinking Machines released Inkling with Apache 2.0 weights, 975B parameters, 41B active parameters, text/image/audio support, and up to 1M context. vLLM, SGLang, Modal, Databricks, and Vercel added day-zero support.
Meta said a model scored 30/30 on the APhO theoretical exam. Team posts described data curation, training and live participation, while public posts questioned which model was evaluated.
Goodside and other testers shared chat links where GPT-5.6 Sol, and sometimes Claude Fable 5, hallucinated hidden messages in noise images or meaningless scribbles. Higher effort settings sometimes did better, but failures reproduced.
Riley Goodside tested random-noise and scribble images with no hidden message. GPT-5.6 Sol often produced invented text, while Claude Fable 5 more often refused or identified the image as non-writing.
LingBot 2.0 released code and weights for a real-time world model and robot-action models. The VLA maps 20 robot body configurations into a 55D action format and filters 90,000 raw robot hours to 50,000 training hours.
Meta launched Muse Spark 1.1 in Meta AI and the Meta Model API public preview for coding, tool use, computer use, and multimodal reasoning. Early eval posts ranked it highly while system-card threads flagged safety details.
Meta launched Muse Image in its apps and previewed Muse Video from the same media-generation family. Meta says Muse Image can reason, search, write code, self-refine, and use test-time compute before generating images.
OpenRouter tested 1,730 visual-reasoning questions across five models and found low-detail images often reduced accuracy while increasing reasoning-token spend. Caps on reasoning effort had the biggest billing impact.
Goodfire introduced Block-Sparse Featurizers, which model activation concepts as multidimensional blocks instead of single SAE directions. The examples cover DINOv3, SDXL, and InceptionV1 activations.
X-Humanoid unveiled TG-VLA as a full-size whole-body VLA framework for humanoids, built around HEX, HAF-VLA, and DSRL-DCT. The company claims DSRL-DCT reached 100% success in mobile-manipulation tasks by freezing the VLA and learning a smaller noise-selection policy.
Gemini Omni Flash ranked #1 on Video Arena at 1404 Elo, 101 points above Seedance 2.0 Mini, and ComfyUI posted a text-prompt video-edit workflow. Google noted the leaderboard is third-party, leaving benchmark provenance as the main caveat.
Google shipped Nano Banana 2 Lite for image generation and Gemini Omni Flash for conversational video generation and editing in the Gemini API and AI Studio. The release sets image generation at about 4 seconds and $0.034 per 1K image, while Omni Flash adds multi-turn video edits at $0.10 per second.
Chandra's developer said Mistral OCR 4 launch numbers for both Chandra and OCR 4 could not be reproduced with public code, and published scripts to show the gaps. The dispute matters because Mistral OCR 4 launched on leaderboard claims, and benchmark settings now directly affect model selection.
Perceptron launched a video_frames input for Mk1 that accepts pre-decoded frames with timestamps instead of forcing clip re-encoding. The change matters for edge and sparse-footage pipelines because 10 minutes of 1080p video can start returning tokens roughly ten times faster.
A day after Seedance 2.0's 4K rollout story, partners began shipping the cheaper Seedance 2.0 Mini across Venice, ComfyUI, and Pika MCP. The 15-second 720p variant with native audio gives video workflows a lower-cost path than the flagship model.
Seedance 2.0 rolled out native 4K generation while Seedance 2.0 Mini landed on fal, Replicate, Pika MCP, and ComfyUI. That matters because engineers can now reach the same video model family through APIs, MCP workflows, and local graph tooling instead of a single web surface.
Baidu released Unlimited OCR as an open-source long-document OCR model with 3B total parameters and 500M active at inference. Early ParseBench testing says it is strong on tables and reading order but weaker on semantic formatting and charts, giving teams a new open-weight OCR option with clear tradeoffs.
OpenRouter released a dedicated Image API that normalizes request shapes across 30-plus models from eight providers. Agents can inspect limits, passthrough options, streaming, and exact per-call cost without hardcoding vendor quirks.
Mistral OCR 4 adds layout-aware extraction with bounding boxes, block typing, and inline confidence across 170 languages. Use it through the API or self-hosted deployments when document pipelines need structure, citations, redaction, and chunking.
Perceptron’s Files API lets developers upload an image or video once and reference it by ID across later requests instead of resending base64 or URLs. That simplifies repeated multimodal workflows and cuts transfer overhead for video-heavy pipelines.
Google put the Interactions API into GA as the new default for Gemini, adding background execution, managed agents, remote sandboxes, and multimodal tools. Builders now get one stateful interface for models, long-running jobs, and future Gemini Omni support.
lift-pdf released an open-source 9B model for schema-constrained document extraction, with code, pip install, playground access, and a 90.2% score on the team's 225-document bench. It matters because the model claims near-Gemini 3.5 Flash accuracy at 9.5s p50, though coverage is still skewed toward Latin-language docs and commercial-use limits remain.
Moonshot rolled out HighSpeed for Kimi K2.7 Code, claiming about 180 tok/s on coding tasks, up to 260 tok/s on shorter contexts, and roughly 6x speedups. Watch the tight capacity limits and mixed benchmark results, and budget for the 2x pricing if you want the faster mode.
ElevenLabs launched Music v2 on ElevenAPI with track generation, reference matching, inpainting, and multilingual output. It gives developers a priced API for commercial music creation and section-level editing.
MiniMax published M3 weights on Hugging Face with 428B total parameters, 23B active parameters, 1M context, and multimodal support. Unsloth quickly added local GGUF builds, so teams can try 2-bit runs at 138GB RAM or VRAM and 3-bit at 165GB.
Zyphra released ZONOS2 under Apache 2.0 with 8B total parameters, 900M active, zero-shot voice cloning, 44.1 kHz DAC audio, and ZTTS1-Eval. The release includes open weights, inference code, and eval code, so teams can run real-time multilingual TTS without a hosted-only stack.
Google released Gemini 3.5 Live Translate for low-latency speech translation across 70+ languages in the Gemini Live API, AI Studio, and Google Translate. The same model is also heading to Google Meet in private preview for Workspace customers.
Posts from WWDC say Apple Intelligence now combines Apple Foundation and Gemini models, and Siri gains visual, on-screen, and app-level actions. Watch for the beta rollout later this year; multiple posts say it will not ship in the EU at launch.
Google released Gemma 4 12B, an Apache 2.0 encoder-free multimodal model with native audio and vision for 16GB-class laptops. Day-zero support in llama.cpp, vLLM, Ollama, MLX, and SGLang should make local agents and on-device apps easier to deploy immediately.
Two days after Qwen 3.7 Plus launched, Hyper, OpenCode, Kilo, and Vals shipped support or rankings around the 1M-context multimodal model. The rapid pickup shows Alibaba’s new model landing quickly in coding-agent tools and public eval stacks outside its own platform.