inference-engine
Projects tagged with inference-engine on GitHub.
12 projects
ds4
CDeepSeek 4 Flash and PRO local inference engine for Metal, CUDA and ROCm
tokenspeed
PythonTokenSpeed is a speed-of-light LLM inference engine.
open-webui
PythonUser-friendly AI Interface (Supports Ollama, OpenAI API, ...)
MTPLX
Python3x decode TPS increase On Qwen 3.6 27B @ temp 0.6 | Native MTP Speculative Decoding On Apple Silicon With No External Drafter.
shard
PythonPipeline-parallel LLM inference across GPUs on separate machines.
inference-school
SwiftA hands-on Swift and Metal course for building LLM inference from first principles on Apple silicon, with 48 guided lessons, runnable exercises, a native macOS Studio, and a complete companion book.
llama.cpp
C++LLM inference in C/C++
cliare
RustCLI agent-readiness measurement, command-shape inference, and CI scorecards
transformers
Python🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.
Audar-ASR-V1
PythonArabic-first generative speech recognition — Audar-ASR-V1 (Flash + Turbo). #1 on the Open Universal Arabic ASR Leaderboard. Model cards, benchmarks & inference.
ollama
GoGet up and running with Kimi-K2.6, GLM-5.2, MiniMax, DeepSeek, gpt-oss, Qwen, Gemma and other models.
local-llm
ShellEverything I know about running LLMs locally