Foundation Model Research
Stories, products, and related signals connected to this tag in Explore.
Stories
Filter storiesAnthropic says an unreleased Claude did not solve the Riemann hypothesis but improved a related zeta-zero lower bound from 41.6% to about 67.2%. Posts describe subagents, expert prompting, and Lean formalization.
Google DeepMind introduced SkillSmith, a method that treats prefix weights or KV-cache states as an input modality so a frozen Gemma 3 4B model can synthesize new skill prefixes at inference time. Reported Composite-SNI Elo improved when cache composition was combined with text descriptions, making it a research artifact rather than a deployable runtime.
OpenAI says an internal Astra model generated arguments for ten long-standing math and theoretical CS problems, with Lean 4 certificates in openai/ten-proofs. Posts focused on the reported sub-$2,000 inference cost.
Meta said a model scored 30/30 on the APhO theoretical exam. Team posts described data curation, training and live participation, while public posts questioned which model was evaluated.
Fable-assisted JLens logs and a separate experiment tested how the proposed J-space forms and whether attention gradients into past tokens can be blocked. The reported blocking method worked in the setup but hurt performance, making it a research update rather than an engineering control.
OpenAI posts said GPT-5.6 Sol helped post-train GPT-5.6 Luna, framing Sol as a research agent rather than just a coding model. Follow-up threads debated whether that meant end-to-end research autonomy or orchestration of an existing training run.
Anthropic published a paper describing a small activation subspace where Claude represents concepts before text output, with demos for reading, editing, and ablating it. Researchers debated whether the method is novel or overstated.