Robotics
Stories, products, and related signals connected to this tag in Explore.
Stories
Filter storiesPerceptron released its 35B Mk1.5 model for drones, quadrupeds, smart glasses, and tool-using agents. The company says the model adds audio, egocentric-video understanding, tracking, and structured outputs, with execution up to 5x faster than Mk1.
Black Forest Labs released FLUX 3 Action, an open 7B model that jointly predicts future video and actions for robot policies. The company reports first place on RoboLab and released embodiment fine-tunes, training recipes, and Jetson deployment support.
Figure reports that Helix 2.5 raised zero-shot household-task success from 9% to 56% across 30 unseen rental homes without retraining. The company released four hours of video showing the humanoid performing the tasks.
Odyssey-3 is a preview world model intended to transfer visual knowledge across robots, vehicles, drones, games, and simulations. Odyssey says downstream control policies need only a few hours of task-specific experience.
World Labs says Atlas combines visual generation with scene reconstruction. It can reconstruct scenes from images, generate camera-controlled frames, and reframe video.
Anthropic opened a research preview of its Model Hardware Standard, a common interface for agents to discover and operate laboratory and manufacturing equipment. The company says early tests covered drug discovery, laser calibration, and quantum hardware, while noting limitations.
Perceptron released weights, inference code, and training details for Isaac 0.5, an embodied model for video perception, reasoning, and robot control. The 36B dynamic-MoE model was trained on 1 million hours of video.
Figure says Index has collected 16 million robot-training video uploads from contributors in 108 countries. The company reports 264,000 app downloads and says contributors upload more than 30 minutes of video each second.
Meta published Muse Spark 1.2 evaluations covering tool-based web page and game creation, robotics planning, and audio-visual tasks. A separate result places it first on Design Arena’s video-to-website benchmark.
Google released Gemini Robotics 2, ER 2, and On-Device 2 for humanoid control, embodied reasoning, and on-device adaptation. Demos showed sub-second streaming and multi-robot task handoffs.
OpenBMB open-sourced MiniCPM-RobotManip, MiniCPM-RobotTrack and PhyAI, claiming local robot tracking, robot memory and throughput gains from 10 Hz to 33-36 Hz. The release packages model artifacts and a runtime path for local robot perception and manipulation experiments.
LingBot 2.0 released code and weights for a real-time world model and robot-action models. The VLA maps 20 robot body configurations into a 55D action format and filters 90,000 raw robot hours to 50,000 training hours.
X-Humanoid unveiled TG-VLA as a full-size whole-body VLA framework for humanoids, built around HEX, HAF-VLA, and DSRL-DCT. The company claims DSRL-DCT reached 100% success in mobile-manipulation tasks by freezing the VLA and learning a smaller noise-selection policy.