⚙️
LLM Inference & Serving
148 tools
Inference runtimes, model serving platforms, fine-tuning infra, and GPU/accelerator providers for LLMs.
OpenRouter
OpenRouter, Inc.
The Unified Interface For LLMs
51 stories
vLLM
vLLM Project
Easy, fast, and cost-efficient LLM serving for everyone.
32 stories
SGLang
LMSYS Corp.
High-performance serving framework for large language and multimodal models.
25 stories
Google AI Studio
Google
Go from prompt to production with Gemini, Veo, Nano Banana, and more
15 stories
Hugging Face
Hugging Face, Inc.
The AI community building the future.
14 stories
Ollama
Ollama Inc.
Build with open models, on your computer and in the cloud
14 stories
llama.cpp
Georgi Gerganov
LLM inference in C/C++
10 stories
O
OpenAI Platform
OpenAI
Build leading AI products on OpenAI’s platform
9 stories
Grok Build
xAI
Bring Grok into your terminal.
7 stories
Together AI
Together AI
The AI Native Cloud
7 stories
Unsloth
Unsloth AI
Easily run & train models locally.
7 stories
Amazon Bedrock
Amazon Web Services
The platform for building generative AI applications and agents at production scale
5 stories
fal
fal
Generative media platform for developers.
4 stories
Wafer
Wafer
The fastest inference on any silicon
4 stories
Baseten
Baseten
Inference Platform: Deploy AI models in production
3 stories
Claude
Anthropic
Meet your thinking partner
3 stories
DFlash
Z Lab
Block Diffusion for Flash Speculative Decoding
3 stories
AI Gateway
Vercel
The AI Gateway for developers
2 stories
Atomic Chat
AtomicMail Systems OÜ
Run local AI models on your device
2 stories
Gemini Live API
Google
Low-latency, real-time voice and vision interactions with Gemini.
2 stories
N
Nous Portal
Nous Research
Everything to power your Hermes Agent
2 stories
NVIDIA NIM
NVIDIA
Designed for rapid, reliable deployment of accelerated generative AI inference anywhere.
2 stories
Arctic RL
Snowflake Inc.
A Unified, Open Source RL Backend for Enterprise Post-Training
1 story
Baidu Qianfan
Baidu, Inc.
An enterprise-grade, all-in-one large model platform centered on Agent development
1 story
Claude Console
Anthropic
Develop with Claude
1 story
Claude Platform on AWS
Anthropic
Access Claude's full platform capabilities through AWS with Anthropic-managed infrastructure.
1 story
Diffusers
Hugging Face
State-of-the-art diffusion models for image, video, and audio generation in PyTorch.
1 story
DSpark
DeepSeek
Confidence-Scheduled Speculative Decoding with Semi-Autoregressive Generation
1 story
Factory Router
Factory
Automatically select the best model for each task.
1 story
Fireworks AI
Fireworks.ai, Inc.
Fast inference and fine-tuning for open source models
1 story
FlashQLA
Alibaba Cloud
High-performance linear attention kernel library built on TileLang
1 story
LM Studio
Element Labs, Inc.
Discover, download, and run local LLMs
1 story
Meta Model API
Meta
Build with Muse Spark on Meta Model API - now in public preview for US developers.
1 story
Miles
RadixArk
Enterprise-Grade Reinforcement Learning for Large-Scale Model Training
1 story
Morph
AutoInfra, Inc.
Inference Built for Coding Agents
1 story
NVIDIA RTX Spark
NVIDIA
A new beginning.
1 story
O
OpenAI Guaranteed Capacity
OpenAI
Guarantee long-term access to OpenAI compute for the products, agents, and customer workflows that matter most.
1 story
PhyAI
MingTi Technology
Latency-first serving engine for Physical AI
1 story
Prime Intellect
Prime Intellect, Inc.
The Open Superintelligence Stack
1 story
Ship
Thesean AI
Save 50% on cost with the same capability and behavior.
1 story
Tile Kernels
Hangzhou DeepSeek Artificial Intelligence Co., Ltd.
Optimized GPU kernels for LLM operations, built with TileLang.
1 story
vLLM-Omni
vLLM Project
Easy, fast, and cheap omni-modality model serving for everyone
1 story
Zyphra Cloud
Zyphra Technologies, Inc.
A full-stack platform for open superintelligence.
1 story
Zyphra Inference
Zyphra
Inference built for agentic intelligence.
1 story
Adaptive Continual Intelligence
Vareon
Continual learning and adaptation after deployment.
0 stories
AFM Playground
Arcee AI
AFM model playground
0 stories
Agent Bricks
Databricks
Production AI agents
0 stories
AI/ML API
AIMLAPI OÜ
One API for 1000+ AI models
0 stories
Arcee Platform
Arcee AI
Building Open Intelligence
0 stories
A
Atlas 950 SuperPoD
Huawei
A SuperPoD AI infrastructure option for large-scale AI training and inference.
0 stories
Baidu AI Studio LLM API
Baidu, Inc.
LLM API for Baidu AI Studio
0 stories
BytePlus
BytePlus Pte. Ltd.
AI-Native Cloud for Enterprise Growth
0 stories
Cerebras Inference
Cerebras Systems
Get instant access to inference that’s up to 15x faster than NVIDIA GPUs.
0 stories
Claude Code Router
musistudio
Manage every agent and provider from one place.
0 stories
ClinePass
Cline Bot Inc.
Best Subscription for Open Weight Models.
0 stories
Cloudflare AI
Cloudflare
Run AI models on Cloudflare's global network.
0 stories
CloudMatrix-Infer
Huawei Cloud
Comprehensive LLM serving solution for Huawei CloudMatrix384
0 stories
Coding Agents
Baseten
The best coding agents run on Baseten
0 stories
Conway
Conway Research
Software product by Conway Research
0 stories
Core AI PyTorch Extensions
Apple
Bring PyTorch models to Core AI for on-device execution.
0 stories
CUDA
NVIDIA
Platform for Accelerated Computing
0 stories
Databricks
Databricks
Data and AI for all
0 stories
DeepAdapt
Vareon
Ship AI that keeps getting better.
0 stories
DeepClaude
Asterisk
Harness the power of DeepSeek R1's reasoning and Claude's creativity and code generation capabilities with a unified API and chat interface.
0 stories
DeepEP
DeepSeek
DeepEP: an efficient expert-parallel communication library
0 stories
DeepGEMM
DeepSeek
clean and efficient BLAS kernel library on GPU
0 stories
dolphin-summarize
Quixi AI
Model architecture summarizer for AI/ML model files
0 stories
DwarfStar
antirez
DeepSeek 4 Flash and PRO local inference engine for Metal, CUDA and ROCm
0 stories
exo
EXO Labs
Run frontier AI locally.
0 stories
FINO
Meta
FIne tuning with NO labels
0 stories
FlashLib
FlashML-org
Bringing Flash Magic to Classical Machine Learning Operators
0 stories
FlashMLA
DeepSeek
Efficient Multi-head Latent Attention Kernels
0 stories
FlashPack
fal
Lightning-Fast Model Loading for PyTorch
0 stories
ForgeTrain
OpenBMB
An LLM Pretraining Framework Built End-to-End by an Autonomous Agent Loop
0 stories
franken_ocr
Jeff Emanuel
Pure-Rust, CPU-only OCR engine with a five-model zoo and no Python, CUDA, ML framework, or GPU required at inference.
0 stories
Gemini Enterprise Agent Platform
Google Cloud
Innovate, build, and deploy enterprise ready agents
0 stories
Gemini Omni
Google
Speak it. See it. Share it.
0 stories
GLM Coding Plan
Z.AI
AI Coding Powered by GLM-5.2 & GLM-5-Turbo for Agents & IDEs
0 stories
Google AI Edge Gallery
Google LLC
Explore, Experience, and Evaluate the Future of On-Device Generative AI with Google AI Edge.
0 stories
GPT4All
Nomic AI
Private & Local AI Chatbot
0 stories
GroqCloud
Groq
Fast LLM inference, OpenAI-compatible. Simple to integrate, easy to scale.
0 stories
gstack
Garry Tan
Use Garry Tan's exact Claude Code setup: 23 opinionated tools that serve as CEO, Designer, Eng Manager, Release Manager, Doc Engineer, and QA
0 stories
GuideLLM
Red Hat
SLO-aware Benchmarking and Evaluation Platform for Optimizing Real-World LLM Inference
0 stories
Inceptron
Inceptron AB
The platform for scalable, reliable, and efficient inference
0 stories
Interfaze
JigsawStack, Inc.
A model built for deterministic tasks
0 stories
Kilo Gateway
Kilo Code Inc.
The Universal Gateway for AI Inference
0 stories
LangSmith LLM Gateway
LangChain
Control every model call
0 stories
Lightning AI
Lightning AI
Idea to AI product, ⚡️ fast.
0 stories
LiteLLM
BerriAI
Open-Source AI Gateway & LLM Proxy
0 stories
llmster
Element Labs, Inc.
LM Studio’s headless daemon
0 stories
LM Link
Element Labs, Inc.
Use your local models, remotely.
0 stories
LMCache
Tensormesh
Supercharge Your LLM with the Fastest KV Cache Layer
0 stories
Mirage
Crisp
Augment your apps with AI
0 stories
Mistral Studio
Mistral AI
Your AI production platform.
0 stories
Modellix
Aurora Mobile Limited
One API for All AI Media Generation
0 stories
ModelScope
Alibaba Cloud
Model-as-a-Service
0 stories
Modular
Modular
Inference reimagined, from Kernel to Cloud. One unified stack.
0 stories
Mooncake
KVCache.AI
A KVCache-centric Disaggregated Architecture for LLM Serving.
0 stories
Multimodal Max
Modular
Multimodal AI software product
0 stories
Naïve
Relixir, Inc.
Ship Apps. Agents. Companies. One prompt. One config file. All your infrastructure.
0 stories
New API
QuantumNous
Next-Generation LLM Gateway and AI Asset Management System
0 stories
Novita Sandbox
Novita AI
Runtime for AI agents to execute real-world tasks
0 stories
NVIDIA DGX B200
NVIDIA Corporation
The foundation for your AI factory.
0 stories
NVIDIA DGX Spark
NVIDIA
Designed to build and run autonomous agents.
0 stories
NVIDIA NeMo RL
NVIDIA
A Scalable and Efficient Post-Training Library
0 stories
N
NVIDIA NVFP4
NVIDIA
4-bit floating-point format and low-precision recipe for efficient AI training and inference.
0 stories
Open Generative AI
MuAPI
Free AI Image & Video Studio
0 stories
Open Instruct
Ai2
AllenAI's post-training codebase
0 stories
O
Open Responses
Open Responses
Open-source specification and ecosystem for multi-provider, interoperable LLM interfaces.
0 stories
OpenMythos
Kye Gomez
Open-source theoretical reconstruction of the Claude Mythos Recurrent-Depth Transformer architecture
0 stories
OptiLLM
Algorithmic SuperIntelligence Labs
2-10x accuracy improvements on reasoning tasks with zero training
0 stories
OrcaRouter
Continuum AI
One AI gateway: adaptive LLM routing & governance
0 stories
Ostris Cloud
Ostris, LLC
AI Toolkit
0 stories
PaddlePaddle AI Studio
Baidu, Inc.
One-stop AI teaching and training platform
0 stories
parakeet.cpp
Frikallo
Fast speech recognition with NVIDIA's Parakeet models in pure C++.
0 stories
Pioneer
Fastino Inc.
Model Routing, Adaptive Inference & Fine-Tuning
0 stories
Platform for AI (PAI)
Alibaba Cloud
Machine Learning Platform for AI
0 stories
Pocket TTS
Kyutai
A TTS that fits in your CPU (and pocket)
0 stories
Prime Intellect Lab
Prime Intellect
Prime Intellect's open research platform for post-training
0 stories
PrismML
PrismML
PrismML platform
0 stories
ProgramAsWeights
ProgramAsWeights
Define functions in English. Run them locally.
0 stories
QuixiCore CUDA
Quixi AI
Tile primitives for speedy kernels
0 stories
Ray Serve LLM
Anyscale
High-performance, scalable framework for deploying Large Language Models (LLMs) in production.
0 stories
Reasonix
esengine
DeepSeek-native coding agent for your terminal
0 stories
R
Rosalind Biodefense
OpenAI
Advancing biological preparedness with trusted developers and government partners.
0 stories
Runpod
Runpod, Inc.
The AI Developer Cloud
0 stories
SiMa.ai
SiMa.ai
Edge AI platform
0 stories
SophontAI
Sophont
Open Medical Superintelligence
0 stories
sparkrun
Spark Arena
Launch, manage, and stop LLM inference workloads on one or more NVIDIA DGX Spark systems — no Slurm, no Kubernetes, no fuss.
0 stories
Step Plan
StepFun
From Coding to Agents, Build It All
0 stories
Still
Baseten
Amortized KV Cache Compaction in a Single Forward Pass
0 stories
SubQ
Subquadratic
The first model built for long-context tasks
0 stories
SubQ
Subquadratic
The first model built for long-context tasks
0 stories
Telnyx
Telnyx
Carrier-owned global communications platform for voice AI agents, SIP trunking, programmable voice, SMS/MMS, and AI inference.
0 stories
Tensordyne Napier
Tensordyne
AI Inference at the speed you want and the margin you need
0 stories
Tenstorrent
Tenstorrent
Compute for every scale.
0 stories
Thinking Machines
Thinking Machines Lab
Thinking Machines software platform
0 stories
TileLang
Tile-AI
A concise domain-specific language designed to streamline the development of high-performance GPU/CPU kernels.
0 stories
TileLang-Ascend
Tile-AI
Ascend TileLang adapter
0 stories
Together Fine-Tuning
Together AI
Fine-tune open-source models for real production use
0 stories
TrainEngine.ai
TrainEngine.ai
Train Dreambooth models. Generate unlimited AI assets.
0 stories
U
Unsloth Studio
Unsloth
Easily run & train models locally.
0 stories
Vast.ai
Vast.ai
GPU cloud marketplace for renting and hosting GPU compute
0 stories
Venice API
Venice.ai
The Only Private AI API
0 stories
vLLM Factory
Latence AI
Production inference for encoders, poolers, and structured prediction — as vLLM plugins.
0 stories
workers-ai-provider
Cloudflare
Workers AI provider for the AI SDK.
0 stories
ZenMux
AI Force Singapore Pte. Ltd.
Unified API for 100+ AI Models
0 stories
阿里云百炼
Alibaba Cloud
一站式大模型开发与应用平台
0 stories