Skip to content
L

llama.cpp

C/C++ inference engine for large language models, supporting quantization

Launch
#27in AI
–In stores

Description

C/C++ inference engine for large language models, supporting quantization, multi-GPU, Apple Silicon, and an OpenAI-compatible server across…

Preview

llama.cpp screenshot 1

Information

People also visit

Apps people often open alongside llama.cpp.

See all
O

Ollama

Load and run large LLMs locally to use in your terminal or build your apps

L

LocalAI

Run LLMs, speech, image generation

A

AutoGPT

Platform for building, deploying, and running AI agents

C

Crawl4AI

Open-source web crawler and scraper

J

Jan

Jan runs open-source AI models on your own hardware or connects

L

LibreChat

LibreChat is a free and open-source chat interface for assistant AIs

llama.cpp FAQ

What is llama.cpp?

C/C++ inference engine for large language models, supporting quantization, multi-GPU, Apple Silicon, and an OpenAI-compatible server across a wide range of hardware.

Do I need to download llama.cpp?

No. llama.cpp runs in your browser at llama.app, so there's nothing to install.

What are the best llama.cpp alternatives?

Popular alternatives to llama.cpp include Ollama, LocalAI and AutoGPT. See all llama.cpp alternatives