⚙️
LLM Inference & Serving
148 tools
Inference runtimes, model serving platforms, fine-tuning infra, and GPU/accelerator providers for LLMs.
Grok Build
xAI
Bring Grok to your computer.
12 stories
Google AI Studio
Google
The fastest way to build with Gemini.
8 stories
fal
fal
The generative media platform for developers
7 stories
Gemini Omni
Google
Speak it. See it. Share it.
4 stories
Ollama
Ollama, Inc.
Get up and running with large language models.
4 stories
BytePlus AI
BytePlus Pte. Ltd.
AI models and products built for real work
3 stories
Claude
Anthropic
Meet Claude, your thinking partner
3 stories
Hugging Face Hub
Hugging Face, Inc.
The AI community building the future.
3 stories
OpenRouter
OpenRouter, Inc.
A unified interface for LLMs.
2 stories
Alibaba Cloud Model Studio
Alibaba Cloud
A one-stop generative AI development platform.
1 story
Claude Console
Anthropic
Build on the Claude Platform
1 story
Gemini Live API
Google
Build real-time, natural conversations with Gemini.
1 story
LiteLLM
BerriAI
One API to call all LLM APIs
1 story
LM Studio
Element Labs, Inc.
Discover, download, and run local LLMs.
1 story
NVIDIA RTX Spark
NVIDIA Corporation
A 128GB unified-memory AI PC platform with 1 PFLOP of FP4 compute
1 story
Adaptive Continual Intelligence
Vareon
Continual Learning After Deployment
0 stories
AFM Playground
Arcee AI
AFM model playground
0 stories
Agent Bricks
Databricks
AI agents for your enterprise
0 stories
AI Gateway
Vercel
One API for all AI models.
0 stories
AI/ML API
AIMLAPI OÜ
One API. 300+ AI Models. All the power you need.
0 stories
Amazon Bedrock
Amazon Web Services
The easiest way to build and scale generative AI applications with foundation models.
0 stories
Arcee Platform
Arcee AI
Enterprise platform for customizing and deploying efficient language models.
0 stories
Arctic RL
Snowflake Inc.
Open-source RL training with ZoRRo acceleration
0 stories
A
Atlas 950 SuperPoD
Huawei
AI computing SuperPoD for large-scale AI training and high-concurrency inference
0 stories
Atomic Chat
AtomicMail Systems OÜ
Local AI app and inference engine for agents.
0 stories
Baidu AI Cloud Qianfan
Baidu, Inc.
AI-native application development platform
0 stories
Baidu AI Studio
Baidu, Inc.
一站式人工智能开发平台
0 stories
Baidu AI Studio
Baidu, Inc.
LLM API for Baidu AI Studio
0 stories
Baseten
Baseten
The inference platform for AI products.
0 stories
Cerebras Inference
Cerebras Systems
The fastest AI inference
0 stories
Claude Code Router
MusiStudio
Use Claude Code as the unified interface for all your LLMs.
0 stories
Claude Platform on AWS
Anthropic
Claude on AWS with native billing, IAM, and Managed Agents
0 stories
ClinePass
Cline Bot Inc.
Best Subscription for Open Weight Models.
0 stories
Cloudflare AI
Cloudflare, Inc.
Build and deploy AI applications on Cloudflare's global network.
0 stories
Cloudflare Workers AI Provider
Cloudflare
Cloudflare Workers AI provider for the AI SDK.
0 stories
CloudMatrix-Infer
Huawei
A comprehensive LLM serving solution
0 stories
Coding Agents
Baseten
The best coding agents run on Baseten
0 stories
Conway Automaton
Conway Research
The first AI that can earn its own existence, replicate, and evolve — without needing a human.
0 stories
Core AI PyTorch Extensions
Apple
Bring PyTorch models to Core AI for on-device execution.
0 stories
CUDA
NVIDIA
CUDA is NVIDIA's parallel computing platform and programming model.
0 stories
Databricks Platform
Databricks
The Data Intelligence Platform
0 stories
DeepAdapt
Vareon Inc.
Turn Deployed AI into Runtime Intelligence
0 stories
DeepClaude
Asterisk
DeepSeek R1 + Claude = DeepClaude
0 stories
DeepEP
DeepSeek
A high-performance expert-parallel communication library.
0 stories
DeepGEMM
DeepSeek
Clean and Efficient FP8 GEMM Library
0 stories
DFlash
Z Lab
Diffusion-based speculative decoding for faster LLM inference
0 stories
Diffusers
Hugging Face
State-of-the-art diffusion models for image, video, and audio generation in PyTorch and FLAX.
0 stories
dolphin-summarize
Quixi AI
A tool for condensed AI/ML model architecture summaries.
0 stories
DSpark
DeepSeek
Confidence-Scheduled Speculative Decoding with Semi-Autoregressive Generation
0 stories
DwarfStar 4
antirez
DeepSeek 4 Flash and PRO local inference engine for Metal, CUDA and ROCm
0 stories
Exo
Exo Labs
Run your own AI cluster at home with everyday devices.
0 stories
Factory Router
Factory
Automatically select the best model for each task.
0 stories
Fine-Tuning
Together AI
Fine-tune open source models.
0 stories
FINO
Meta
Metadata-driven adaptation of vision foundation models for scientific domains.
0 stories
Fireworks AI
Fireworks.ai, Inc.
The fastest way to build with generative AI
0 stories
FlashLib
FlashML-org
Bringing Flash Magic to Classical Machine Learning Operators
0 stories
FlashMLA
DeepSeek
Efficient MLA decoding kernels for Hopper GPUs.
0 stories
FlashPack
fal
High-throughput tensor loading for PyTorch -- load models at up to 25Gbps without GDS.
0 stories
FlashQLA
Alibaba Cloud
High-performance linear attention kernels built on TileLang.
0 stories
ForgeTrain
OpenBMB
A fully AI-generated pretraining framework.
0 stories
FrankenOCR
Jeff Emanuel
Pure Rust local OCR for agents
0 stories
Gemini Enterprise
Google Cloud
The AI platform for work.
0 stories
GLM Coding Plan
Z.ai
Coding-focused access to GLM models
0 stories
Google AI Edge Gallery
Google LLC
Run AI models locally on your Android device.
0 stories
GPT4All
Nomic AI
Chat with local LLMs, privately.
0 stories
GroqCloud
Groq
Fast AI inference in the cloud
0 stories
gstack
Garry Tan
Superpowers for Claude Code
0 stories
GuideLLM
Red Hat
Benchmark LLM inference performance
0 stories
Inceptron
Inceptron AB
The platform for scalable, reliable, and efficient inference
0 stories
Interfaze
JigsawStack, Inc.
A model built for deterministic tasks
0 stories
Kilo Gateway
Kilo Code Inc.
The Universal Gateway for AI Inference
0 stories
Lab
Prime Intellect
The full-stack platform for training your own models
0 stories
LangSmith LLM Gateway
LangChain
Control every model call
0 stories
Lighthouse Attention
Nous Research
Long Context Pre-Training with Lighthouse Attention
0 stories
Lightning AI Platform
Lightning AI
Idea to AI product, ⚡️ fast.
0 stories
llama.cpp
GGML.ai
LLM inference in C/C++
0 stories
llmster
Element Labs, Inc.
LM Studio’s headless daemon
0 stories
LM Link
Element Labs, Inc.
Use your local models, remotely.
0 stories
LMCache
TensorMesh
Supercharge your LLM with KV Cache.
0 stories
MAX
Modular
Multimodal AI software product
0 stories
Meta Model API
Meta
Meta Model API public preview
0 stories
Miles
RadixArk
An asynchronous reinforcement-learning framework for large language models
0 stories
Mirage
Crisp
Augment your apps with AI
0 stories
Mistral AI Studio
Mistral AI
Build your frontier with Studio.
0 stories
Modellix
Aurora Mobile Limited
One AI API, flattened parameters.
0 stories
ModelScope
Alibaba Cloud
ModelScope: discover and use AI models, datasets, and applications.
0 stories
Modular Platform
Modular
AI infrastructure that runs everywhere.
0 stories
Mooncake
KVCache.AI
A KVCache-centric disaggregated architecture for LLM serving.
0 stories
Morph
AutoInfra, Inc.
Fast inference for coding agents
0 stories
Naïve
Relixir, Inc.
Ship Apps. Agents. Companies. One prompt. One config file. All your infrastructure.
0 stories
NeMo RL
NVIDIA
Scalable reinforcement learning for LLM post-training.
0 stories
New API
QuantumNous
AI Model Aggregation Management and Distribution System
0 stories
N
Nous Portal
Nous Research
Everything to power your Hermes Agent
0 stories
Novita Sandbox
Novita AI
Runtime for AI agents to execute real-world tasks
0 stories
NVFP4
NVIDIA
Efficient and accurate low-precision inference.
0 stories
NVIDIA DGX B200
NVIDIA Corporation
The AI supercomputer for every enterprise.
0 stories
NVIDIA DGX Spark
NVIDIA
Desktop AI supercomputer for AI developers.
0 stories
NVIDIA NIM
NVIDIA
AI inference microservices
0 stories
Open Generative AI
MuAPI
Free AI Image & Video Studio
0 stories
Open Instruct
Allen Institute for AI
A fully open-source framework for instruction tuning and preference tuning of LMs.
0 stories
O
Open Responses
Open Responses
Open-source specification and ecosystem for building multi-provider, interoperable LLM interfaces.
0 stories
OpenAI Guaranteed Capacity
OpenAI
Guarantee long-term access to OpenAI compute.
0 stories
O
OpenAI Platform
OpenAI
Build with OpenAI
0 stories
OpenMythos
The Swarm Corporation
Open-source theoretical reconstruction of the Claude Mythos Recurrent-Depth Transformer architecture
0 stories
OptiLLM
Algorithmic SuperIntelligence Labs
Optimizing LLMs.
0 stories
OrcaRouter
Continuum AI
One gateway. Every model.
0 stories
Ostris Cloud
Ostris, LLC
AI Toolkit
0 stories
parakeet.cpp
Frikallo
Fast speech recognition with NVIDIA's Parakeet models in pure C++.
0 stories
PhyAI
OpenBMB
Enabling robots to understand, remember, and act.
0 stories
Pioneer
Fastino Labs
Everything your inference stack is missing.
0 stories
Platform for AI (PAI)
Alibaba Cloud
Machine Learning Platform for AI
0 stories
Pocket TTS
Kyutai
A lightweight text-to-speech model.
0 stories
Prime Intellect Platform
Prime Intellect, Inc.
Open infrastructure for AI research and superintelligence.
0 stories
PrismML
PrismML
PrismML platform
0 stories
ProgramAsWeights
ProgramAsWeights
Define functions in English. Run them locally.
0 stories
QuixiCore CUDA
Quixi AI
Native NVIDIA CUDA kernels for Ampere and newer GPUs.
0 stories
Ray Serve LLM
Anyscale
Build fast, scalable, and cost-effective LLM services.
0 stories
Reasonix
esengine
A coding agent you can leave running.
0 stories
R
Rosalind Biodefense Program
OpenAI
Advancing biological preparedness with trusted developers and government partners.
0 stories
RunPod
RunPod, Inc.
The Cloud Built for AI
0 stories
SGLang
LMSYS Org
A fast serving framework for large language models and vision language models.
0 stories
Ship
Thesean AI
Inference optimization endpoint targeting lower-cost execution while preserving reference-model behavior.
0 stories
SiMa.ai
SiMa.ai
Edge AI platform
0 stories
SophontAI
Sophont
Open Medical Superintelligence
0 stories
sparkrun
Spark Arena
Launch, manage, and stop LLM inference workloads on one or more NVIDIA DGX Spark systems — no Slurm, no Kubernetes, no fuss.
0 stories
Step Plan
StepFun
From Coding to Agents, Build It All
0 stories
Still
Baseten
Amortized KV Cache Compaction in a Single Forward Pass
0 stories
Telnyx
Telnyx LLC
The full-stack communications platform.
0 stories
Tensordyne Napier
Tensordyne
AI inference at the speed you want and the margin you need
0 stories
Tenstorrent
Tenstorrent
Compute for every scale.
0 stories
Thinking Machines
Thinking Machines Lab
Thinking Machines software platform
0 stories
TileKernels
DeepSeek
Open-source TileLang kernels for DeepSeek workloads
0 stories
TileLang
Tile-AI
A domain-specific language designed to streamline the development of high-performance GPU kernels for AI workloads.
0 stories
TileLang-Ascend
Tile-AI
Ascend TileLang adapter
0 stories
Together AI Platform
Together AI
The AI Acceleration Cloud
0 stories
TrainEngine.ai
TrainEngine.ai
Train Dreambooth models. Generate unlimited AI assets.
0 stories
UncommonRoute
Commonstack
Automatic model routing for lower LLM spend.
0 stories
Unsloth
Unsloth AI
Fine-tune LLMs and VLMs faster with less memory.
0 stories
Unsloth Studio
Unsloth
Run and train AI models locally with Unsloth Studio.
0 stories
Vast.ai
Vast.ai Inc.
The world's largest GPU marketplace.
0 stories
Venice API
Venice AI
Private, unrestricted access to all the leading AI models across text, image, video, and audio, behind one API key.
0 stories
vLLM
vLLM Project
A high-throughput and memory-efficient inference and serving engine for LLMs.
0 stories
vLLM Factory
Latence AI
Production inference for encoders, poolers, and structured prediction — as vLLM plugins.
0 stories
vLLM-Omni
vLLM Project
High-performance serving framework for omni-modal models.
0 stories
Wafer
Wafer, Inc.
Fast AI inference for open models
0 stories
ZenMux
AI Force Singapore Pte. Ltd.
One API for all AI models.
0 stories
Zyphra Cloud
Zyphra Technologies, Inc.
A full-stack platform for open superintelligence.
0 stories
Zyphra Inference
Zyphra
Serverless inference for frontier open-weight models focused on long horizon agentic workloads.
0 stories