dorkhub

awesome-local-llm

A curated list of awesome platforms, tools, practices and resources that helps run LLMs locally

rafska
2.5k305 forksMITupdated 1 week ago
git clone https://github.com/rafska/awesome-local-llm.gitrafska/awesome-local-llm

Awesome local LLM

A curated list of awesome platforms, tools, practices and resources that helps run LLMs locally

Table of Contents

Inference platforms

  • LM Studio - discover, download and run local LLMs
  • unsloth - unified web UI for training and running open models like Qwen, DeepSeek, and Gemma locally
  • LocalAI - the free, open-source alternative to OpenAI, Claude and others
  • jan - an open source alternative to ChatGPT that runs 100% offline on your computer
  • ChatBox - user-friendly desktop client app for AI models/LLMs
  • lemonade - a local LLM server with GPU and NPU Acceleration

Back to Table of Contents

Inference engines

  • ollama - get up and running with LLMs
  • llama.cpp - LLM inference in C/C++
  • vllm - a high-throughput and memory-efficient inference and serving engine for LLMs
  • exo - run your own AI cluster at home with everyday devices
  • BitNet - official inference framework for 1-bit LLMs
  • sglang - a fast serving framework for large language models and vision language models
  • TensorRT-LLM - provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs
  • Nano-vLLM - a lightweight vLLM implementation built from scratch
  • omlx - LLM inference server with continuous batching & SSD caching for Apple Silicon — managed from the macOS menu bar
  • koboldcpp - run GGUF models easily with a KoboldAI UI
  • mistral.rs - fast, flexible LLM inference
  • dynamo - a datacenter scale distributed inference serving framework
  • flashinfer - kernel library for LLM serving
  • mlx-lm - generate text and fine-tune large language models on Apple silicon with MLX
  • gpustack - simple, scalable AI model deployment on GPU clusters
  • LiteRT-LM - Google's production-ready, high-performance, open-source inference framework for deploying Large Language Models on edge devices
  • mlx-vlm - a package for inference and fine-tuning of Vision Language Models (VLMs) on your Mac using MLX
  • executorch - on-device AI across mobile, embedded and edge for PyTorch
  • mini-sglang - a lightweight yet high-performance inference framework for Large Language Models
  • distributed-llama - connect home devices into a powerful cluster to accelerate LLM inference
  • LiteRT - Google's on-device framework for high-performance ML & GenAI deployment on edge platforms, via efficient conversion, runtime, and optimization
  • ik_llama.cpp - llama.cpp fork with additional SOTA quants and improved performance
  • aphrodite-engine - large-scale LLM inference engine
  • FastFlowLM - run LLMs on AMD Ryzen™ AI NPUs
  • tokenspeed - a speed-of-light LLM inference engine
  • krasis - a Hybrid LLM runtime which focuses on efficient running of larger models on consumer grade VRAM limited hardware
  • vllm-gfx906 - vLLM for AMD gfx906 GPUs, e.g. Radeon VII / MI50 / MI60
  • llm-scaler - run LLMs on Intel Arc™ Pro B60 GPUs

Back to Table of Contents

User Interfaces

  • Open WebUI - User-friendly AI Interface (Supports Ollama, OpenAI API, ...)
  • Lobe Chat - an open-source, modern design AI chat framework
  • Text generation web UI - LLM UI with advanced features, easy setup, and multiple backend support
  • SillyTavern - LLM Frontend for Power Users
  • Page Assist - Use your locally running AI models to assist you in your web browsing

Back to Table of Contents

Large Language Models

Explorers, Benchmarks, Leaderboards

  • Arena - benchmark & compare the best AI models
  • AI Models & API Providers Analysis - understand the AI landscape to choose the best model and provider for your use case
  • SWE-rebench - a continuously evolving and decontaminated benchmark for software engineering LLMs
  • BullshitBench - measure whether AI models challenge nonsensical prompts instead of confidently answering them
  • LLM Explorer - explore list of the open-source LLM models
  • Dubesor LLM Benchmark table - small-scale manual performance comparison benchmark
  • oobabooga benchmark - a list sorted by size (on disk) for each score
  • CyberGym - evaluating AI agents' real-world cybersecurity capabilities at scale
  • vakra - a benchmark for evaluating multi-hop, multi-source tool-calling in AI agents

Back to Table of Contents

Model providers

  • Qwen - powered by Alibaba Cloud
  • Mistral AI - a pioneering French artificial intelligence startup
  • Tencent - a profile of a Chinese multinational technology conglomerate and holding company
  • Unsloth AI - focusing on making AI more accessible to everyone (GGUFs etc.)
  • bartowski - providing GGUF versions of popular LLMs
  • Beijing Academy of Artificial Intelligence - a private non-profit organization engaged in AI research and development
  • Open Thoughts - a team of researchers and engineers curating the best open reasoning datasets

Back to Table of Contents

Specific models

General purpose

  • Qwen3.6 - a collection of the latest generation Qwen LLMs
  • NVIDIA Nemotron v3 - a family of open models from NVIDIA with open weights, training data and recipes, delivering leading efficiency and accuracy for building specialized AI agents
  • Gemma 4 - a family of open models built by Google DeepMind, that are multimodal, handling text and image input (with audio supported on small models) and generating text output
  • Mistral Medium 3.5 - The first flaship models from Mistral AI handling instruction-following, reasoning, and coding in a single set of opened-weights
  • gpt-oss - a collection of open-weight models from OpenAI, designed for powerful reasoning, agentic tasks, and versatile developer use cases
  • gpt-oss-puzzle-88B - a deployment-optimized large language model developed by NVIDIA, derived from OpenAI's gpt-oss-120b
  • Hunyuan - a collection of Tencent's open-source efficient LLMs designed for versatile deployment across diverse computational environments
  • Phi-4 - a family of small language, multi-modal and reasoning models from Microsoft
  • OpenReasoning-Nemotron - a collection of models from NVIDIA, trained on 5M reasoning traces for math, code and science
  • Kimi K2.5 - a collection of open-source, native multimodal agentic models from Moonshot AI that advances practical capabilities in long-horizon coding, coding-driven design, proactive autonomous execution, and swarm-based task orchestration
  • GLM-5.2 - a Z.ai's flagship model for long-horizon tasks
  • Granite 4.1 - efficient language models from IBM for multilingual generation, coding, RAG, and AI assistant workflows
  • EXAONE-4.5 - LG's First Open-Weight Vision-Language Model for Industrial Intelligence
  • ERNIE 4.5 - a collection of large-scale multimodal models from Baidu
  • Seed-OSS - a collection of LLMs developed by ByteDance's Seed Team, designed for powerful long-context, reasoning, agent and general capabilities, and versatile developer-friendly features
  • Step-3.5-Flash - most capable open-source foundation model, engineered to deliver frontier reasoning and agentic capabilities with exceptional efficiency
  • Rio-3.5-Open-397B - a frontier-class general-purpose AI model post-trained from Qwen 3.5 397B
  • Nex-N2 - a collection of agent models built for real-world productivity scenarios

Back to Table of Contents

Coding

  • Qwen3-Coder-Next - a collection of Qwen's open-weight language models designed specifically for coding agents and local development
  • Devstral 2 - a couple of agentic LLMs for software engineering tasks, excelling at using tools to explore codebases, edit multiple files, and power SWE Agents
  • Mellum 2 - an assistant model trained by JetBrain
  • MiniMax-M3 - a native multimodal model with 1M context
  • MiniMax-M2 - a collection of SOTA models for real-world dev & agents
  • Laguna-S-2.1 - a 118B total parameter Mixture-of-Experts model with 8B activated parameters per token, designed for agentic coding and long-horizon work
  • SWE-FastContext - a family of code-search models from Microsoft powering the Explore subagent for coding agents
  • OmniCoder-9B - a 9-billion parameter coding agent model built by Tesslate, fine-tuned on top of Qwen3.5-9B's hybrid architecture
  • NousCoder-14B - a competitive programming model post-trained on Qwen3-14B via reinforcement learning
  • MusaCoder-27B - a code model developed by Moore Threads for PyTorch-to-CUDA/MUSA native kernel generation

Back to Table of Contents

Multimodal

  • Qwen3-Omni - a collection of the natively end-to-end multilingual omni-modal foundation models from Qwen
  • GLM-4.6V - a collection of open source multimodal models with native tool use from Zhipu AI

Back to Table of Contents

Image

  • Qwen-Image - a collection of models for image generation, edit and decomposition from Qwen
  • Qwen3-VL - a collection of the most powerful vision-language models in the Qwen series to date
  • GLM-Image - an image generation model
  • Granite Vision - multimodal models from IBM built for visual document analysis and image understanding
  • HunyuanImage - a collection of image generation models from Tencent
  • HunyuanVideo - a collection of video generation models from Tencent
  • Vidi - a collection of models for multimodal video understanding and creation
  • FastVLM - a collection of VLMs with efficient vision encoding from Apple
  • MiniCPM-o & MiniCPM-V - multimodal models with leading performance
  • LFM2-VL - a colection of vision-language models, designed for on-device deployment
  • ClipTagger-12b - a vision-language model (VLM) designed for video understanding at massive scale

Back to Table of Contents

Audio

  • whisper-large-v3 - a state-of-the-art model for automatic speech recognition (ASR) and speech translation from OpenAI
  • Nemotron Speech - a collection of open, state-of-the-art, production‑ready enterprise speech models from the NVIDIA Speech research team for ASR, TTS, Speaker Diarization and S2S
  • Qwen3-ASR - a collection of models that support language identification and ASR for 52 languages and dialects
  • Qwen3-TTS - a collection of TTS models that cover 10 major languages as well as multiple dialectal voice profiles to meet global application needs
  • Granite Speech - a collection of compact and efficient speech-language models from IBM, specifically designed for multilingual automatic speech recognition (ASR) and bidirectional automatic speech translation (AST)
  • Voxtral-Small-24B-2507 - an enhancement of Mistral Small 3, incorporating state-of-the-art audio input capabilities while retaining best-in-class text performance
  • Voxtral-Mini-4B-Realtime-2602 - a multilingual, realtime speech-transcription model and among the first open-source solutions to achieve accuracy comparable to offline systems with a delay of <500ms
  • Voxtral-4B-TTS-2603 - frontier, open-weights text-to-speech model that’s fast, instantly adaptable, and produces lifelike speech for voice agents
  • chatterbox - first production-grade open-source TTS model
  • VibeVoice - a collection of frontier text-to-speech models from Microsoft
  • Kitten TTS - a collection of open-source realistic text-to-speech models designed for lightweight deployment and high-quality voice synthesis
  • Streaming Sortformer Diarizer 4spk v2.1 - a streaming version of a novel end-to-end neural model for speaker diarization from NVIDIA

Back to Table of Contents

Retrieval-Augmented Generation

  • Nemotron RAG - a set of tools to build retrieval-augmented generation (RAG) systems, improve search and ranking accuracy, and extract structured data from complex docs
  • Qwen3-Embedding - a collection of the latest proprietary Qwen models, specifically designed for text embedding and ranking tasks
  • Qwen3-VL-Embedding - an addition to the Qwen embedding models, specifically designed for multimodal information retrieval and cross-modal understanding
  • Qwen3-Reranker - a collection of the latest proprietary Qwen models, engineered to refine embedding results
  • Qwen3-VL-Reranker - an addition to the Qwen embedding models, specifically designed for multimodal information retrieval and cross-modal understanding

Back to Table of Contents

Safeguards

  • Granite Guardian - a collection of safety models from IBM for detecting risks, toxicity, and hallucinations in LLM workflows
  • Qwen3Guard - a collection of safety moderation models built upon Qwen3
  • NemoGuard - a collection of models from NVIDIA for content safety, topic-following and security guardrails
  • Nemotron-3.5-Content-Safety - a small language model (SLM) that uses Google's Gemma-3-4B-it as the base and is fine-tuned by NVIDIA on multimodal, multilingual, and reasoning-oriented content-safety datasets
  • Privasis - a collection of lightweight text-sanitization models from NVIDIA designed to remove or abstract sensitive information from text according to a user-provided sanitization instruction
  • SingGuard - a collection of policy-adaptive multimodal LLM Guardrails with dynamic reasoning
  • HARC - a family of safety-aligned instruction models from Microsoft trained with HARC
  • gpt-oss-safeguard - a collection of safety reasoning models built-upon gpt-oss from OpenAI
  • privacy-filter - a bidirectional token-classification model from OpenAI for personally identifiable information (PII) detection and masking in text
  • AprielGuard - a safeguard model designed to detect and mitigate both safety risks and security threats in LLM interactions

Back to Table of Contents

Miscellaneous

  • Marco-MoE - a suit of multilingual MoE models with highly-sparse architectures
  • Jan-v3 - a 4B baseline model for fine-tuning, designed for downstream work: improved instruction following out of the box, strong starting point for fine-tuning and effective lightweight coding assistance
  • Jan-v2-VL - a family of VLM focused on reliable, many-step task execution
  • Nemotron-Orchestrator-8B - a state-of-the-art 8B orchestration model designed to solve complex, multi-turn agentic tasks by coordinating a diverse set of expert models and tools
  • Arch-Router-1.5B - the fastest LLM router model that aligns to subjective usage preferences
  • Waypoint - a collection of real-time interactive video world models
  • Hunyuan3D - a collection of everything related (models, datasets etc.) to 3D assets generation from Tencent
  • Hunyuan-GameCraft-1.0 - a novel framework for high-dynamic interactive video generation in game environments
  • void-model - a model from Netflix that removes objects from videos along with all interactions they induce on the scene — not just secondary effects like shadows and reflections, but physical interactions like objects falling when a person is removed

Back to Table of Contents

Tools

Models

  • llmfit - hundreds of models & providers, one command to find what runs on your hardware
  • outlines - structured outputs for LLMs
  • llama-swap - reliable model swapping for any local OpenAI compatible server - llama.cpp, vllm, etc.
  • llguidance - super-fast structured outputs

Back to Table of Contents

Agent Frameworks

  • AutoGPT - a powerful platform that allows you to create, deploy, and manage continuous AI agents that automate complex workflows
  • langflow - a powerful tool for building and deploying AI-powered agents and workflows
  • langchain - build context-aware reasoning applications
  • anything-llm - the all-in-one Desktop & Docker AI application with built-in RAG, AI agents, No-code agent builder, MCP compatibility, and more
  • autogen - a programming framework for agentic AI
  • Flowise - build AI agents, visually
  • pi - AI agent toolkit: coding agent CLI, unified LLM API, TUI & web UI libraries, Slack bot, vLLM pods
  • llama_index - the leading framework for building LLM-powered agents over your data
  • crewAI - a framework for orchestrating role-playing, autonomous AI agents
  • agno - a full-stack framework for building Multi-Agent Systems with memory, knowledge and reasoning
  • sim - open-source platform to build and deploy AI agent workflows
  • openai-agents-python - a lightweight, powerful framework for multi-agent workflows
  • NemoClaw - run OpenClaw more securely inside NVIDIA OpenShell with managed inference
  • SuperAGI - an open-source framework to build, manage and run useful Autonomous AI Agents
  • camel - the first and the best multi-agent framework
  • pydantic-ai - a Python agent framework designed to help you quickly, confidently, and painlessly build production grade applications and workflows with Generative AI
  • txtai - all-in-one open-source AI framework for semantic search, LLM orchestration and language model workflows
  • agent-framework - a framework for building, orchestrating and deploying AI agents and multi-agent workflows with support for Python and .NET
  • archgw - a high-performance proxy server that handles the low-level work in building agents: like applying guardrails, routing prompts to the right agent, and unifying access to LLMs, etc.
  • genkit - open-source framework for building AI-powered apps in JavaScript, Go, and Python, built and used in production by Google
  • ClaraVerse - privacy-first, fully local AI workspace with Ollama LLM chat, tool calling, agent builder, Stable Diffusion, and embedded n8n-style automation
  • NeMo-Agent-Toolkit - an open-source library for efficiently connecting and optimizing teams of AI agents
  • ragbits - building blocks for rapid development of GenAI applications

Back to Table of Contents

Model Context Protocol

  • mindsdb - federated query engine for AI - the only MCP Server you'll ever need
  • github-mcp-server - GitHub's official MCP Server
  • playwright-mcp - Playwright MCP server
  • chrome-devtools-mcp - Chrome DevTools for coding agents
  • n8n-mcp - a MCP for Claude Desktop / Claude Code / Windsurf / Cursor to build n8n workflows for you
  • awslabs/mcp - AWS MCP Servers — helping you get the most out of AWS, wherever you use MCP
  • mcp-atlassian - MCP server for Atlassian tools (Confluence, Jira)
  • dbhub - zero-dependency, token-efficient database MCP server for Postgres, MySQL, SQL Server, MariaDB, SQLite

Back to Table of Contents

Retrieval-Augmented Generation

  • pathway - Python ETL framework for stream processing, real-time analytics, LLM pipelines and RAG
  • graphrag - a modular graph-based RAG system
  • LightRAG - simple and fast RAG
  • haystack - AI orchestration framework to build customizable, production-ready LLM applications, best suited for building RAG, question answering, semantic search or conversational agent chatbots
  • vanna - an open-source Python RAG framework for SQL generation and related functionality
  • graphiti - build real-time knowledge graphs for AI Agents
  • onyx - the AI platform connected to your company's docs, apps, and people
  • claude-context - make entire codebase the context for any coding agent
  • pipeshub-ai - a fully extensible and explainable workplace AI platform for enterprise search and workflow automation

Back to Table of Contents

Coding Agents

  • opencode - a AI coding agent built for the terminal
  • zed - a next-generation code editor designed for high-performance collaboration with humans and AI
  • OpenHands - a platform for software development agents powered by AI
  • cline - autonomous coding agent right in your IDE, capable of creating/editing files, executing commands, using the browser, and more with your permission every step of the way
  • aider - AI pair programming in your terminal
  • tabby - an open-source GitHub Copilot alternative, set up your own LLM-powered code completion server
  • continue - create, share, and use custom AI code assistants with our open-source IDE extensions and hub of models, rules, prompts, docs, and other building blocks
  • void - an open-source Cursor alternative, use AI agents on your codebase, checkpoint and visualize changes, and bring any model or host locally
  • goose - an open-source, extensible AI agent that goes beyond code suggestions
  • Roo-Code - a whole dev team of AI agents in your code editor
  • crush - the glamourous AI coding agent for your favourite terminal
  • kilocode - open source AI coding assistant for planning, building, and fixing code
  • humanlayer - the best way to get AI coding agents to solve hard problems in complex codebases
  • 99 - neovim AI agent done right
  • ProxyAI - the leading open-source AI copilot for JetBrains

Back to Table of Contents

Computer Use

  • open-interpreter - a natural language interface for computers
  • OmniParser - a simple screen parsing tool towards pure vision based GUI agent
  • openwork - an open-source alternative to Claude Cowork, powered by OpenCode
  • cua - the Docker Container for Computer-Use AI Agents
  • Agent-S - an open agentic framework that uses computers like a human
  • self-operating-computer - a framework to enable multimodal models to operate a computer
  • OpenRoom - a browser-based desktop where AI Agent operates every app through natural language, from MiniMaxAI

Back to Table of Contents

Browser Automation

  • puppeteer - a JavaScript API for Chrome and Firefox
  • playwright - a framework for Web Testing and Automation
  • browser-use - make websites accessible for AI agents
  • firecrawl - turn entire websites into LLM-ready markdown or structured data
  • stagehand - the AI Browser Automation Framework
  • nanobrowser - open-source Chrome extension for AI-powered web automation

Back to Table of Contents

Memory Management

  • mem0 - universal memory layer for AI Agents
  • mempalace - the highest-scoring AI memory system ever benchmarked
  • letta - the stateful agents framework with memory, reasoning, and context management
  • supermemory - memory engine and app that is extremely fast, scalable
  • cognee - memory for AI Agents in 5 lines of code
  • LMCache - supercharge your LLM with the fastest KV Cache Layer
  • memU - an open-source memory framework for AI companions
  • reasoning-bank - a memory mechanism for agents that learns from both successful and failed trajectories, with reasoning stored as memory content

Back to Table of Contents

Testing, Evaluation and Observability

  • langfuse - an open-source LLM engineering platform: LLM Observability, metrics, evals, prompt management, playground, datasets. Integrates with OpenTelemetry, Langchain, OpenAI SDK, LiteLLM, and more
  • opik - debug, evaluate, and monitor your LLM applications, RAG systems, and agentic workflows with comprehensive tracing, automated evaluations, and production-ready dashboards
  • openllmetry - an open-source observability for your LLM application, based on OpenTelemetry
  • giskard - an open-source evaluation & testing for AI & LLM systems
  • agenta - an open-source LLMOps platform: prompt playground, prompt management, LLM evaluation, and LLM observability all in one place
  • Evaluator - open-source library for scalable, reproducible evaluation of AI models and benchmarks

Back to Table of Contents

Research

  • Perplexica - an open-source alternative to Perplexity AI, the AI-powered search engine
  • gpt-researcher - an LLM based autonomous agent that conducts deep local and web research on any topic and generates a long report with citations
  • SurfSense - an open-source alternative to NotebookLM / Perplexity / Glean
  • open-notebook - an open-source implementation of Notebook LM with more flexibility and features
  • RD-Agent - automate the most critical and valuable aspects of the industrial R&D process
  • local-deep-researcher - fully local web research and report writing assistant
  • local-deep-research - an AI-powered research assistant for deep, iterative research
  • maestro - an AI-powered research application designed to streamline complex research tasks

Back to Table of Contents

Training and Fine-tuning

  • heretic - fully automatic censorship removal for language models
  • sentence-transformers - a Python library for using and training embedding and reranker models for applications like retrieval augmented generation, semantic search, and more
  • trl - train transformer language models with reinforcement learning
  • OpenRLHF - an easy-to-use, high-performance open-source RLHF framework built on Ray, vLLM, ZeRO-3 and HuggingFace Transformers, designed to make RLHF training simple and accessible
  • slime - an LLM post-training framework for RL Scaling
  • Kiln - the easiest tool for fine-tuning LLM models, synthetic data generation, and collaborating on datasets
  • OpenEnv - an interface library for RL post training with environments
  • augmentoolkit - train an open-source LLM on new facts

more like this

perplexity-cli

🧠 A simple command-line client for the Perplexity API. Ask questions and receive answers directly from the terminal! 🚀🚀🚀

Python176

AeroPath

:hugs: AeroPath: An airway segmentation benchmark dataset with challenging pathology

Jupyter Notebook53

BrainAI

BrainAI is a set of helper classes to add AI to your game.

C#53

search

search projects, people, and tags