Training Infrastructure
Stories, products, and related signals connected to this tag in Explore.
Stories
Filter storiesLangSmith Fine-Tuning and the open-source smithtune CLI now turn agent traces into fine-tuning datasets. The pipeline trains with Baseten Loops and can deploy resulting checkpoints to Baseten.
Xiaomi released MiMo V2.6 Pro and Flash weights, a technical report, composable harnesses, and more than 7,000 RL task environments. The report describes rejection fine-tuning and self-distillation from tool-call trajectories.
MiMo is reportedly livestreaming RL training for its V2.6 Pro and Flash models, publishing batch data, harness composition, reward curves, and infrastructure metrics. Reported cost figures list the trillion-parameter Pro run at about $493,000.
Magic says a new pretraining recipe matched DeepSeek V4 Pro with roughly 50 times less compute. After a 10x scale-up costing about $4 million, the company says it exceeded publicly available base models.
A practitioner reports using GPT-6 Astra Ultra to collect data, generate masks, and train RF-DETR-Seg-M models without human-supplied labels. Other demos show instance-segmentation and video-frame annotation workflows.
Meta says AIRA3 placed eighth among roughly 4,000 teams in a live NVIDIA Kaggle competition. The system used many long-running agents to fine-tune a 30B Nemotron model on a private test set.
Open Athena has begun training Marin, a 535B-parameter MoE with 23B active parameters, over 18T tokens. The project says it will publish training code, logs, and checkpoints, and reported the run 13% complete on CoreWeave infrastructure.
Figure says Index has collected 16 million robot-training video uploads from contributors in 108 countries. The company reports 264,000 app downloads and says contributors upload more than 30 minutes of video each second.
Marin has begun an open training run for a 535B-parameter mixture-of-experts model with 23B active parameters. The team plans to train on 18.75T tokens across 11 GB200 NVL72 systems over about three months.
OpenAI paused some deployment-focused frontier reinforcement-learning training to strengthen security and monitoring. Its largest planned frontier RL run remains on hold while the company gathers alignment evidence.
Reports say GLM-5.3 retained GLM-5.2's base model while post-training raised Terminal-Bench from 4.6 to 28.3 and DeepSWE from 46.2 to 66.9. A technical account attributes the gains to RL infrastructure changes.
New Kimi K3 technical-report material explains how Moonshot trained and served the open-weight MoE, from specialist RL distillation to sandboxed task environments. Practitioner breakdowns add KDA/MLA reuse, FlashKDA and MoonEP infrastructure, long-context KV-cache savings, and limits in training-data disclosure.
A German consortium released the small SOOFI sovereign base model trained on 27T tokens. Analysts said it reuses Nemotron 3 Nano architecture and many hyperparameters with a changed data mix, and benchmarks drew criticism for overstating capability versus Qwen and Nemotron.
OpenAI posts said GPT-5.6 Sol helped post-train GPT-5.6 Luna, framing Sol as a research agent rather than just a coding model. Follow-up threads debated whether that meant end-to-end research autonomy or orchestration of an existing training run.