All topics

inference-engine

Projects tagged with inference-engine on GitHub.

12 projects

ds4 logo

ds4

C
90

DeepSeek 4 Flash and PRO local inference engine for Metal, CUDA and ROCm

18.8k+8960Jan 21, 1970
tokenspeed logo

tokenspeed

Python
85

TokenSpeed is a speed-of-light LLM inference engine.

1.6k+650Jan 21, 1970
open-webui logo

open-webui

Python
95

User-friendly AI Interface (Supports Ollama, OpenAI API, ...)

146.8k+9140Jan 21, 1970
MTPLX logo

MTPLX

Python
82

3x decode TPS increase On Qwen 3.6 27B @ temp 0.6 | Native MTP Speculative Decoding On Apple Silicon With No External Drafter.

1.1k+420Jan 21, 1970
shard logo

shard

Python
77

Pipeline-parallel LLM inference across GPUs on separate machines.

432+50Jan 21, 1970
inference-school logo

inference-school

Swift
82

A hands-on Swift and Metal course for building LLM inference from first principles on Apple silicon, with 48 guided lessons, runnable exercises, a native macOS Studio, and a complete companion book.

1770Jan 21, 1970
llama.cpp logo

llama.cpp

C++
90

LLM inference in C/C++

121.9k+1.1k0Jan 21, 1970
cliare logo

cliare

Rust
82

CLI agent-readiness measurement, command-shape inference, and CI scorecards

7120Jan 21, 1970
transformers logo

transformers

Python
90

🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.

163.1k+2660Jan 21, 1970
Audar-ASR-V1 logo

Audar-ASR-V1

Python
73

Arabic-first generative speech recognition — Audar-ASR-V1 (Flash + Turbo). #1 on the Open Universal Arabic ASR Leaderboard. Model cards, benchmarks & inference.

5580Jan 21, 1970
ollama logo

ollama

Go
90

Get up and running with Kimi-K2.6, GLM-5.2, MiniMax, DeepSeek, gpt-oss, Qwen, Gemma and other models.

177.1k+5350Jan 21, 1970
local-llm logo

local-llm

Shell
77

Everything I know about running LLMs locally

1.6k+4230Jan 21, 1970
inference-engine Open Source Projects | MushyBook