⚙️
LLM Inference & Serving
147 tools
Inference runtimes, model serving platforms, fine-tuning infra, and GPU/accelerator providers for LLMs.
AI Studio
Google
Build with Gemini.
6 stories
Grok Build
xAI
Bring Grok into your terminal.
6 stories
Gemini Omni
Google
Google's AI assistant
4 stories
BytePlus
BytePlus Pte. Ltd.
Enterprise technology platform
3 stories
fal
fal
Generative media platform for developers.
3 stories
Ollama
Ollama Inc.
Get up and running with large language models.
3 stories
Claude
Anthropic
Meet your thinking partner
2 stories
Hugging Face
Hugging Face, Inc.
The AI community building the future.
2 stories
OpenRouter
OpenRouter, Inc.
One API for many models.
2 stories
Claude Console
Anthropic
Develop with Claude
1 story
LM Studio
Element Labs, Inc.
Run local language models on your computer.
1 story
NVIDIA RTX Spark
NVIDIA
A new beginning.
1 story
SubQ
Subquadratic
SubQ
1 story
Adaptive Continual Intelligence
Vareon
Continual learning and adaptation after deployment.
0 stories
AFM Playground
Arcee AI
AFM model playground
0 stories
Agent Bricks
Databricks
Production AI agents
0 stories
AI Gateway
Vercel
The AI Gateway for developers
0 stories
AI/ML API
AIMLAPI OÜ
One API for 1000+ AI models
0 stories
Amazon Bedrock
Amazon Web Services
Build and scale generative AI applications with foundation models.
0 stories
Arcee
Arcee AI
Enterprise AI platform for language models
0 stories
Arctic RL
Snowflake Inc.
A Unified, Open Source RL Backend for Enterprise Post-Training
0 stories
A
Atlas 950 SuperPoD
Huawei
A SuperPoD AI infrastructure option for large-scale AI training and inference.
0 stories
Atomic Chat
AtomicMail Systems OÜ
Run local AI models on your device
0 stories
Baidu AI Studio LLM API
Baidu, Inc.
LLM API for Baidu AI Studio
0 stories
Baidu Qianfan
Baidu, Inc.
An enterprise-grade, all-in-one large model platform centered on Agent development
0 stories
Baseten
Baseten
AI inference platform
0 stories
Cerebras
Cerebras Systems
Hosted AI inference and API service
0 stories
Claude Code Router
musistudio
Manage every agent and provider from one place.
0 stories
Claude Platform on AWS
Anthropic
Claude on AWS
0 stories
ClinePass
Cline Bot Inc.
Best Subscription for Open Weight Models.
0 stories
Cloudflare AI Platform
Cloudflare
Build full-stack AI applications on Cloudflare
0 stories
CloudMatrix-Infer
Huawei Cloud
Comprehensive LLM serving solution for Huawei CloudMatrix384
0 stories
Coding Agents
Baseten
Coding Agents
0 stories
Conway
Conway Research
Software product by Conway Research
0 stories
Core AI PyTorch Extensions
Apple
Bring PyTorch models to Core AI for on-device execution.
0 stories
CUDA
NVIDIA
Platform for Accelerated Computing
0 stories
Databricks
Databricks
Data and AI for all
0 stories
DeepAdapt
Vareon
Ship AI that keeps getting better.
0 stories
DeepClaude
Asterisk
Claude-powered reasoning and chat
0 stories
DeepEP
DeepSeek
Efficient expert-parallel communication library
0 stories
DeepGEMM
DeepSeek
Clean and efficient FP8 GEMM kernels with fine-grained scaling.
0 stories
DFlash
Z Lab
Block Diffusion for Flash Speculative Decoding
0 stories
DGX Spark
NVIDIA
Personal AI platform from NVIDIA
0 stories
Diffusers
Hugging Face
State-of-the-art pretrained diffusion models
0 stories
dolphin-summarize
Quixi AI
Model architecture summarizer for AI/ML model files
0 stories
DSpark
DeepSeek
Confidence-Scheduled Speculative Decoding with Semi-Autoregressive Generation
0 stories
DwarfStar
antirez
Software product by antirez
0 stories
exo
EXO Labs
Run frontier AI locally.
0 stories
Factory Router
Factory
Automatically select the best model for each task.
0 stories
FINO
Meta
FIne tuning with NO labels
0 stories
Fireworks AI
Fireworks.ai, Inc.
Fast, scalable infrastructure for generative AI.
0 stories
FlashLib
FlashML-org
Bringing Flash Magic to Classical Machine Learning Operators
0 stories
FlashMLA
DeepSeek
Efficient MLA decoding kernel for Hopper GPUs
0 stories
FlashPack
fal
High-throughput tensor loading for PyTorch -- load models at up to 25Gbps without GDS.
0 stories
FlashQLA
Alibaba Cloud
FlashQLA
0 stories
ForgeTrain
OpenBMB
An LLM Pretraining Framework Built End-to-End by an Autonomous Agent Loop
0 stories
franken_ocr
Jeff Emanuel
Pure-Rust, CPU-only OCR engine with a five-model zoo and no Python, CUDA, ML framework, or GPU required at inference.
0 stories
Gemini Live API
Google
Real-time multimodal API for Gemini.
0 stories
GLM Coding Plan
Z.AI
AI Coding Powered by GLM-5.2 & GLM-5-Turbo for Agents & IDEs
0 stories
Google AI Edge Gallery
Google LLC
Explore and run AI models on-device.
0 stories
GPT4All
Nomic AI
Run local LLMs on your own hardware.
0 stories
Groq
Groq
Fast AI inference platform
0 stories
gstack
Garry Tan
Unverified product associated with Garry Tan.
0 stories
GuideLLM
Red Hat
SLO-aware Benchmarking and Evaluation Platform for Optimizing Real-World LLM Inference
0 stories
Inceptron
Inceptron AB
The platform for scalable, reliable, and efficient inference
0 stories
Interfaze
JigsawStack, Inc.
AI-assisted interface builder
0 stories
Kilo Gateway
Kilo Code Inc.
The Universal Gateway for AI Inference
0 stories
LangSmith LLM Gateway
LangChain
Control every model call
0 stories
Lightning AI
Grid.ai, Inc.
Build, train, and deploy AI apps.
0 stories
LiteLLM
BerriAI
OpenAI-compatible API for all LLMs
0 stories
llama.cpp
Georgi Gerganov
LLM inference in C/C++
0 stories
llmster
Element Labs, Inc.
llmster
0 stories
LM Link
Element Labs, Inc.
Product details unavailable
0 stories
LMCache
Tensormesh
Supercharge Your LLM with the Fastest KV Cache Layer
0 stories
Meta Model API
Meta
Build with Muse Spark on Meta Model API - now in public preview for US developers.
0 stories
Miles
RadixArk
Enterprise-Grade Reinforcement Learning for Large-Scale Model Training
0 stories
Mirage
Crisp
Augment your apps with AI
0 stories
Mistral Studio
Mistral AI
Your AI production platform.
0 stories
ModellixAI
Modellix
AI software product
0 stories
ModelScope
Alibaba Cloud
Model-as-a-Service
0 stories
Modular
Modular
Inference reimagined, from Kernel to Cloud. One unified stack.
0 stories
Mooncake
KVCache.AI
A KVCache-centric Disaggregated Architecture for LLM Serving.
0 stories
Morph
AutoInfra, Inc.
Inference Built for Coding Agents
0 stories
Multimodal Max
Modular
Multimodal AI software product
0 stories
Naïve
Relixir, Inc.
Software product named Naïve
0 stories
NeMo-RL
NVIDIA
A scalable and efficient post-training library
0 stories
New API
QuantumNous
Next-Generation LLM Gateway and AI Asset Management System
0 stories
N
Nous Portal
Nous Research
Everything to power your Hermes Agent
0 stories
Novita Sandbox
Novita AI
Runtime for AI agents to execute real-world tasks
0 stories
NVIDIA DGX B200
NVIDIA Corporation
The foundation for your AI factory.
0 stories
NVIDIA NIM
NVIDIA
NVIDIA NIM microservices
0 stories
N
NVIDIA NVFP4
NVIDIA
4-bit floating-point format and low-precision recipe for efficient AI training and inference.
0 stories
Open Generative AI
MuAPI
Free AI Image & Video Studio
0 stories
Open Instruct
Ai2
AllenAI's post-training codebase
0 stories
O
Open Responses
Open Responses
Open-source specification and ecosystem for multi-provider, interoperable LLM interfaces.
0 stories
O
OpenAI Guaranteed Capacity
OpenAI
Guaranteed access to OpenAI capacity
0 stories
O
OpenAI Platform
OpenAI
Build leading AI products on OpenAI’s platform
0 stories
OpenMythos
Kye Gomez
OpenMythos
0 stories
OptiLLM
Algorithmic SuperIntelligence Labs
2-10x accuracy improvements on reasoning tasks with zero training
0 stories
OrcaRouter
Continuum AI
One AI gateway: adaptive LLM routing & governance
0 stories
Ostris Cloud
Ostris, LLC
AI Toolkit
0 stories
PaddlePaddle AI Studio
Baidu, Inc.
One-stop AI teaching and training platform
0 stories
parakeet.cpp
Frikallo
C++ software project
0 stories
Pioneer
Fastino Inc.
Software product by Fastino Inc.
0 stories
Platform for AI (PAI)
Alibaba Cloud
Machine Learning Platform for AI
0 stories
Pocket TTS
Kyutai
A TTS that fits in your CPU (and pocket)
0 stories
Prime Intellect
Prime Intellect
Decentralized AI platform
0 stories
Prime Intellect Lab
Prime Intellect
Prime Intellect Lab
0 stories
PrismML
PrismML
PrismML platform
0 stories
ProgramAsWeights
ProgramAsWeights
Define functions in English. Run them locally.
0 stories
QuixiCore CUDA
Quixi AI
Tile primitives for speedy kernels
0 stories
Ray Serve LLM
Anyscale
Serving LLMs
0 stories
Reasonix
esengine
DeepSeek-native coding agent for your terminal
0 stories
R
Rosalind Biodefense
OpenAI
Advancing biological preparedness with trusted developers and government partners.
0 stories
Runpod
Runpod, Inc.
The AI Developer Cloud
0 stories
SGLang
LMSYS Corp.
High-performance serving framework for large language and multimodal models.
0 stories
Ship
Thesean AI
Ship reduces your LLM bill 50% without changing model behavior. Guaranteed by an SLA.
0 stories
SiMa.ai
SiMa.ai
Edge AI platform
0 stories
SophontAI
Sophont
Open Medical Superintelligence
0 stories
sparkrun
Spark Arena
Official details could not be verified from source access in this run.
0 stories
Step Plan
StepFun
StepFun software product
0 stories
Still
Baseten
Amortized KV Cache Compaction in a Single Forward Pass
0 stories
Subquadratic
Subquadratic
Official Subquadratic product
0 stories
Telnyx
Telnyx
Carrier-owned global communications platform for voice AI agents, SIP trunking, programmable voice, SMS/MMS, and AI inference.
0 stories
Tensordyne
Tensordyne
AI Math, Re-Engineered. AI Inference, Re-Defined.
0 stories
Tenstorrent
Tenstorrent
Compute for every scale.
0 stories
Thinking Machines
Thinking Machines Lab
Thinking Machines software platform
0 stories
Tile Kernels
Hangzhou DeepSeek Artificial Intelligence Co., Ltd.
Optimized GPU kernels for LLM operations, built with TileLang.
0 stories
TileLang
Tile-AI
A concise domain-specific language designed to streamline the development of high-performance GPU/CPU kernels.
0 stories
TileLang-Ascend
Tile-AI
Ascend TileLang adapter
0 stories
Together AI
Together AI
AI development platform for open models
0 stories
Together Fine-Tuning
Together AI
Fine-tune open-source models for real production use
0 stories
TrainEngine.ai
TrainEngine.ai
Train Dreambooth models. Generate unlimited AI assets.
0 stories
Unsloth
Unsloth AI
Fast LLM fine-tuning
0 stories
U
Unsloth Studio
Unsloth
Unsloth Studio
0 stories
Vast.ai
Vast.ai
GPU cloud marketplace for renting and hosting GPU compute
0 stories
Venice API
Venice.ai
Private, unrestricted access to all the leading AI models across text, image, video, and audio, behind one API key.
0 stories
Vertex AI
Google Cloud
Build, deploy, and scale AI applications with Vertex AI.
0 stories
vLLM
vLLM Project
Easy, fast, and cheap LLM serving for everyone
0 stories
vLLM Factory
Latence AI
Production inference for encoders, poolers, and structured prediction — as vLLM plugins.
0 stories
vLLM-Omni
vLLM Project
Easy, fast, and cheap omni-modality model serving for everyone
0 stories
Wafer
Wafer
The fastest inference on any silicon
0 stories
workers-ai-provider
Cloudflare
Workers AI provider for the AI SDK.
0 stories
ZenMux
AI Force Singapore Pte. Ltd.
Unified API for all models, intelligent routing, and AI Model Insurance.
0 stories
Zyphra Cloud
Zyphra Technologies Inc.
AI cloud from Zyphra
0 stories
Zyphra Inference
Zyphra
Inference for Zyphra models
0 stories
阿里云百炼
Alibaba Cloud
一站式大模型开发及应用构建平台
0 stories