⚙️
LLM Inference & Serving
148 tools
Inference runtimes, model serving platforms, fine-tuning infra, and GPU/accelerator providers for LLMs.
OpenRouter
OpenRouter, Inc.
A unified interface for LLMs.
56 stories
vLLM
vLLM Project
A high-throughput and memory-efficient inference and serving engine for LLMs.
42 stories
SGLang
LMSYS Org
A fast serving framework for large language models and vision language models.
28 stories
Hugging Face Hub
Hugging Face, Inc.
The AI community building the future.
17 stories
Google AI Studio
Google
The fastest way to build with Gemini.
16 stories
Ollama
Ollama, Inc.
Get up and running with large language models.
16 stories
llama.cpp
GGML.ai
LLM inference in C/C++
12 stories
Grok Build
xAI
Bring Grok to your computer.
11 stories
O
OpenAI Platform
OpenAI
Build with OpenAI
10 stories
Together AI Platform
Together AI
The AI Acceleration Cloud
8 stories
Unsloth
Unsloth AI
Fine-tune LLMs and VLMs faster with less memory.
7 stories
Amazon Bedrock
Amazon Web Services
The easiest way to build and scale generative AI applications with foundation models.
5 stories
fal
fal
The generative media platform for developers
5 stories
Claude
Anthropic
Meet Claude, your thinking partner
4 stories
Wafer
Wafer, Inc.
Fast AI inference for open models
4 stories
AI Gateway
Vercel
One API for all AI models.
3 stories
Baseten
Baseten
The inference platform for AI products.
3 stories
DFlash
Z Lab
Diffusion-based speculative decoding for faster LLM inference
3 stories
Meta Model API
Meta
Meta Model API public preview
3 stories
Atomic Chat
AtomicMail Systems OÜ
Local AI app and inference engine for agents.
2 stories
Diffusers
Hugging Face
State-of-the-art diffusion models for image, video, and audio generation in PyTorch and FLAX.
2 stories
DSpark
DeepSeek
Confidence-Scheduled Speculative Decoding with Semi-Autoregressive Generation
2 stories
Gemini Live API
Google
Build real-time, natural conversations with Gemini.
2 stories
LM Studio
Element Labs, Inc.
Discover, download, and run local LLMs.
2 stories
N
Nous Portal
Nous Research
Everything to power your Hermes Agent
2 stories
NVIDIA DGX Spark
NVIDIA
Desktop AI supercomputer for AI developers.
2 stories
NVIDIA NIM
NVIDIA
AI inference microservices
2 stories
vLLM-Omni
vLLM Project
High-performance serving framework for omni-modal models.
2 stories
Arctic RL
Snowflake Inc.
Open-source RL training with ZoRRo acceleration
1 story
Baidu AI Cloud Qianfan
Baidu, Inc.
AI-native application development platform
1 story
Claude Console
Anthropic
Build on the Claude Platform
1 story
Claude Platform on AWS
Anthropic
Claude on AWS with native billing, IAM, and Managed Agents
1 story
Coding Agents
Baseten
The best coding agents run on Baseten
1 story
Factory Router
Factory
Automatically select the best model for each task.
1 story
Fireworks AI
Fireworks.ai, Inc.
The fastest way to build with generative AI
1 story
FlashQLA
Alibaba Cloud
High-performance linear attention kernels built on TileLang.
1 story
Gemini Enterprise
Google Cloud
The AI platform for work.
1 story
Miles
RadixArk
An asynchronous reinforcement-learning framework for large language models
1 story
ModelScope
Alibaba Cloud
ModelScope: discover and use AI models, datasets, and applications.
1 story
Morph
AutoInfra, Inc.
Fast inference for coding agents
1 story
NVIDIA RTX Spark
NVIDIA Corporation
A 128GB unified-memory AI PC platform with 1 PFLOP of FP4 compute
1 story
OpenAI Guaranteed Capacity
OpenAI
Guarantee long-term access to OpenAI compute.
1 story
parakeet.cpp
Frikallo
Fast speech recognition with NVIDIA's Parakeet models in pure C++.
1 story
PhyAI
OpenBMB
Enabling robots to understand, remember, and act.
1 story
Prime Intellect Platform
Prime Intellect, Inc.
Open infrastructure for AI research and superintelligence.
1 story
Ship
Thesean AI
Inference optimization endpoint targeting lower-cost execution while preserving reference-model behavior.
1 story
TileKernels
DeepSeek
Open-source TileLang kernels for DeepSeek workloads
1 story
Zyphra Cloud
Zyphra Technologies, Inc.
A full-stack platform for open superintelligence.
1 story
Zyphra Inference
Zyphra
Serverless inference for frontier open-weight models focused on long horizon agentic workloads.
1 story
Adaptive Continual Intelligence
Vareon
Continual Learning After Deployment
0 stories
AFM Playground
Arcee AI
AFM model playground
0 stories
Agent Bricks
Databricks
AI agents for your enterprise
0 stories
AI/ML API
AIMLAPI OÜ
One API. 300+ AI Models. All the power you need.
0 stories
Alibaba Cloud Model Studio
Alibaba Cloud
A one-stop generative AI development platform.
0 stories
Arcee Platform
Arcee AI
Enterprise platform for customizing and deploying efficient language models.
0 stories
A
Atlas 950 SuperPoD
Huawei
AI computing SuperPoD for large-scale AI training and high-concurrency inference
0 stories
Baidu AI Studio
Baidu, Inc.
一站式人工智能开发平台
0 stories
Baidu AI Studio
Baidu, Inc.
LLM API for Baidu AI Studio
0 stories
BytePlus AI
BytePlus Pte. Ltd.
AI models and products built for real work
0 stories
Cerebras Inference
Cerebras Systems
The fastest AI inference
0 stories
Claude Code Router
MusiStudio
Use Claude Code as the unified interface for all your LLMs.
0 stories
ClinePass
Cline Bot Inc.
Best Subscription for Open Weight Models.
0 stories
Cloudflare AI
Cloudflare, Inc.
Build and deploy AI applications on Cloudflare's global network.
0 stories
Cloudflare Workers AI Provider
Cloudflare
Cloudflare Workers AI provider for the AI SDK.
0 stories
CloudMatrix-Infer
Huawei
A comprehensive LLM serving solution
0 stories
Conway Automaton
Conway Research
The first AI that can earn its own existence, replicate, and evolve — without needing a human.
0 stories
Core AI PyTorch Extensions
Apple
Bring PyTorch models to Core AI for on-device execution.
0 stories
CUDA
NVIDIA
CUDA is NVIDIA's parallel computing platform and programming model.
0 stories
Databricks Platform
Databricks
The Data Intelligence Platform
0 stories
DeepAdapt
Vareon Inc.
Turn Deployed AI into Runtime Intelligence
0 stories
DeepClaude
Asterisk
DeepSeek R1 + Claude = DeepClaude
0 stories
DeepEP
DeepSeek
A high-performance expert-parallel communication library.
0 stories
DeepGEMM
DeepSeek
Clean and Efficient FP8 GEMM Library
0 stories
dolphin-summarize
Quixi AI
A tool for condensed AI/ML model architecture summaries.
0 stories
DwarfStar 4
antirez
DeepSeek 4 Flash and PRO local inference engine for Metal, CUDA and ROCm
0 stories
Exo
Exo Labs
Run your own AI cluster at home with everyday devices.
0 stories
Fine-Tuning
Together AI
Fine-tune open source models.
0 stories
FINO
Meta
Metadata-driven adaptation of vision foundation models for scientific domains.
0 stories
FlashLib
FlashML-org
Bringing Flash Magic to Classical Machine Learning Operators
0 stories
FlashMLA
DeepSeek
Efficient MLA decoding kernels for Hopper GPUs.
0 stories
FlashPack
fal
High-throughput tensor loading for PyTorch -- load models at up to 25Gbps without GDS.
0 stories
ForgeTrain
OpenBMB
A fully AI-generated pretraining framework.
0 stories
FrankenOCR
Jeff Emanuel
Pure Rust local OCR for agents
0 stories
Gemini Omni
Google
Speak it. See it. Share it.
0 stories
GLM Coding Plan
Z.ai
Coding-focused access to GLM models
0 stories
Google AI Edge Gallery
Google LLC
Run AI models locally on your Android device.
0 stories
GPT4All
Nomic AI
Chat with local LLMs, privately.
0 stories
GroqCloud
Groq
Fast AI inference in the cloud
0 stories
gstack
Garry Tan
Superpowers for Claude Code
0 stories
GuideLLM
Red Hat
Benchmark LLM inference performance
0 stories
Inceptron
Inceptron AB
The platform for scalable, reliable, and efficient inference
0 stories
Interfaze
JigsawStack, Inc.
A model built for deterministic tasks
0 stories
Kilo Gateway
Kilo Code Inc.
The Universal Gateway for AI Inference
0 stories
Lab
Prime Intellect
The full-stack platform for training your own models
0 stories
LangSmith LLM Gateway
LangChain
Control every model call
0 stories
Lighthouse Attention
Nous Research
Long Context Pre-Training with Lighthouse Attention
0 stories
Lightning AI Platform
Lightning AI
Idea to AI product, ⚡️ fast.
0 stories
LiteLLM
BerriAI
One API to call all LLM APIs
0 stories
llmster
Element Labs, Inc.
LM Studio’s headless daemon
0 stories
LM Link
Element Labs, Inc.
Use your local models, remotely.
0 stories
LMCache
TensorMesh
Supercharge your LLM with KV Cache.
0 stories
MAX
Modular
Multimodal AI software product
0 stories
Mirage
Crisp
Augment your apps with AI
0 stories
Mistral AI Studio
Mistral AI
Build your frontier with Studio.
0 stories
Modellix
Aurora Mobile Limited
One AI API, flattened parameters.
0 stories
Modular Platform
Modular
AI infrastructure that runs everywhere.
0 stories
Mooncake
KVCache.AI
A KVCache-centric disaggregated architecture for LLM serving.
0 stories
Naïve
Relixir, Inc.
Ship Apps. Agents. Companies. One prompt. One config file. All your infrastructure.
0 stories
NeMo RL
NVIDIA
Scalable reinforcement learning for LLM post-training.
0 stories
New API
QuantumNous
AI Model Aggregation Management and Distribution System
0 stories
Novita Sandbox
Novita AI
Runtime for AI agents to execute real-world tasks
0 stories
NVFP4
NVIDIA
Efficient and accurate low-precision inference.
0 stories
NVIDIA DGX B200
NVIDIA Corporation
The AI supercomputer for every enterprise.
0 stories
Open Generative AI
MuAPI
Free AI Image & Video Studio
0 stories
Open Instruct
Allen Institute for AI
A fully open-source framework for instruction tuning and preference tuning of LMs.
0 stories
O
Open Responses
Open Responses
Open-source specification and ecosystem for building multi-provider, interoperable LLM interfaces.
0 stories
OpenMythos
The Swarm Corporation
Open-source theoretical reconstruction of the Claude Mythos Recurrent-Depth Transformer architecture
0 stories
OptiLLM
Algorithmic SuperIntelligence Labs
Optimizing LLMs.
0 stories
OrcaRouter
Continuum AI
One gateway. Every model.
0 stories
Ostris Cloud
Ostris, LLC
AI Toolkit
0 stories
Pioneer
Fastino Labs
Everything your inference stack is missing.
0 stories
Platform for AI (PAI)
Alibaba Cloud
Machine Learning Platform for AI
0 stories
Pocket TTS
Kyutai
A lightweight text-to-speech model.
0 stories
PrismML
PrismML
PrismML platform
0 stories
ProgramAsWeights
ProgramAsWeights
Define functions in English. Run them locally.
0 stories
QuixiCore CUDA
Quixi AI
Native NVIDIA CUDA kernels for Ampere and newer GPUs.
0 stories
Ray Serve LLM
Anyscale
Build fast, scalable, and cost-effective LLM services.
0 stories
Reasonix
esengine
A coding agent you can leave running.
0 stories
R
Rosalind Biodefense Program
OpenAI
Advancing biological preparedness with trusted developers and government partners.
0 stories
RunPod
RunPod, Inc.
The Cloud Built for AI
0 stories
SiMa.ai
SiMa.ai
Edge AI platform
0 stories
SophontAI
Sophont
Open Medical Superintelligence
0 stories
sparkrun
Spark Arena
Launch, manage, and stop LLM inference workloads on one or more NVIDIA DGX Spark systems — no Slurm, no Kubernetes, no fuss.
0 stories
Step Plan
StepFun
From Coding to Agents, Build It All
0 stories
Still
Baseten
Amortized KV Cache Compaction in a Single Forward Pass
0 stories
Telnyx
Telnyx LLC
The full-stack communications platform.
0 stories
Tensordyne Napier
Tensordyne
AI inference at the speed you want and the margin you need
0 stories
Tenstorrent
Tenstorrent
Compute for every scale.
0 stories
Thinking Machines
Thinking Machines Lab
Thinking Machines software platform
0 stories
TileLang
Tile-AI
A domain-specific language designed to streamline the development of high-performance GPU kernels for AI workloads.
0 stories
TileLang-Ascend
Tile-AI
Ascend TileLang adapter
0 stories
TrainEngine.ai
TrainEngine.ai
Train Dreambooth models. Generate unlimited AI assets.
0 stories
UncommonRoute
Commonstack
Automatic model routing for lower LLM spend.
0 stories
Unsloth Studio
Unsloth
Run and train AI models locally with Unsloth Studio.
0 stories
Vast.ai
Vast.ai Inc.
The world's largest GPU marketplace.
0 stories
Venice API
Venice AI
Private, unrestricted access to all the leading AI models across text, image, video, and audio, behind one API key.
0 stories
vLLM Factory
Latence AI
Production inference for encoders, poolers, and structured prediction — as vLLM plugins.
0 stories
ZenMux
AI Force Singapore Pte. Ltd.
One API for all AI models.
0 stories