Tabby: Self-Hosted GitHub Copilot Alternative
Want to go deeper than this article?
Free account unlocks the first chapter of all 25 courses — RAG, agents, MCP, voice AI, MLOps, real GitHub repos.
Picked your coding model? Build a real AI dev workflow. From local copilots to agents that ship code — the structured path, running on your hardware. First chapter free.
Published on April 10, 2026 -- 19 min read
GitHub Copilot costs $10/month per developer and sends every keystroke to Microsoft's servers. For a 10-person team, that is $1,200/year -- and your proprietary code passes through infrastructure you do not control. Tabby eliminates both problems: it is free, open-source, and runs entirely on your network.
Tabby has 23,000+ GitHub stars, supports VS Code, JetBrains, Vim, and Neovim, and serves real-time code completions from models you choose. The architecture is deliberately small: one server process holding one model in GPU memory, an HTTP endpoint, and an editor extension that calls it. That is why the install is short and the failure modes are few.
Set your expectations by the model, not by the server. Tabby is plumbing — the suggestions you get are whatever a 3B to 7B open code model produces. Those models are strong at the completions that make up most of autocomplete's value (finishing a line, filling a known pattern, matching the shape of the code around the cursor) and weaker at long, novel blocks, which is where a frontier cloud model still pulls ahead.
Here is how to set it up, which model to pick, and how to tune it for your team.
What is Tabby
Tabby is an open-source AI code completion server built by TabbyML. It provides:
- Real-time inline completion served over HTTP from a model held in local GPU memory
- Multi-IDE support: VS Code, JetBrains (IntelliJ, PyCharm, WebStorm), Vim, Neovim
- Model flexibility: StarCoder2, DeepSeek-Coder, CodeLlama, Qwen2.5-Coder
- Repository indexing: learns your codebase patterns for better suggestions
- Admin dashboard: user management, usage analytics, model configuration
- Enterprise features: LDAP/OAuth auth, audit logging, access controls
The architecture is simple: Tabby runs as a server (standalone binary, Docker container, or Homebrew install), loads a code model into GPU memory, and serves completions over HTTP. IDE extensions connect to the server and inject suggestions inline, identical to how Copilot works.
What Tabby Does Not Do
Tabby focuses specifically on code completion -- the autocomplete experience. It does not include:
- Chat interface (use Continue.dev or Claude for that)
- Code explanation or documentation generation
- Agent mode or autonomous task execution
- Code review or PR analysis
This focused scope is actually an advantage: Tabby does one thing and does it well, with minimal resource usage.
Reading articles is good. Building is better.
Free account = the first chapter of all 25 courses, with a per-chapter AI tutor. No card.
Why Self-Host Code Completion
Privacy
Every character you type in Copilot gets sent to GitHub's servers. For companies handling regulated data (healthcare, finance, defense), customer PII, or proprietary algorithms, that is a compliance problem. Self-hosted Tabby keeps all code on your network.
Cost at Scale
| Team Size | Copilot Cost/Year | Tabby Hardware Cost | Tabby Breakeven |
|---|---|---|---|
| 5 devs | $600 | $1,600 (RTX 4090) | 2.7 years |
| 10 devs | $1,200 | $1,600 (RTX 4090) | 1.3 years |
| 25 devs | $3,000 | $1,600 (RTX 4090) | 6.4 months |
| 50 devs | $6,000 | $4,500 (A6000 48GB) | 9 months |
| 100 devs | $12,000 | $4,500 (A6000 48GB) | 4.5 months |
At 10+ developers, Tabby pays for itself in the first year. At 25+, the savings are substantial.
Customization
Copilot gives you one model, take it or leave it. Tabby lets you:
- Choose models optimized for your languages (DeepSeek-Coder for Python, StarCoder2 for polyglot)
- Index your private repositories for context-aware completions
- Fine-tune models on your codebase (advanced)
- Control context window size and completion behavior
Uptime Independence
Copilot goes down when GitHub has an outage. Your Tabby server runs on your infrastructure, on your schedule.
Installation Methods
Method 1: Docker (Recommended for Teams)
Docker is the cleanest way to deploy Tabby, especially for team servers.
# NVIDIA GPU (CUDA)
docker run -it \
--gpus all \
-p 8080:8080 \
-v $HOME/.tabby:/data \
tabbyml/tabby \
serve --model StarCoder2-3B --device cuda
# AMD GPU (ROCm)
docker run -it \
--device /dev/kfd --device /dev/dri \
--group-add video \
-p 8080:8080 \
-v $HOME/.tabby:/data \
tabbyml/tabby-rocm \
serve --model StarCoder2-3B --device rocm
After startup, open http://localhost:8080 in your browser. You will see the admin dashboard where you can create user accounts, manage models, and view usage analytics.
Method 2: Homebrew (macOS)
# Install
brew install tabbyml/tabby/tabby
# Run with Apple Metal acceleration
tabby serve --model StarCoder2-3B --device metal
# Verify it is running
curl http://localhost:8080/v1/health
Method 3: Direct Binary (Linux)
# Download the latest release
curl -L https://github.com/TabbyML/tabby/releases/latest/download/tabby_x86_64-unknown-linux-gnu -o tabby
chmod +x tabby
# Run with CUDA
./tabby serve --model StarCoder2-3B --device cuda
# Or run as a systemd service for persistence
sudo tee /etc/systemd/system/tabby.service << 'EOF'
[Unit]
Description=Tabby AI Code Completion Server
After=network.target
[Service]
Type=simple
User=tabby
ExecStart=/usr/local/bin/tabby serve --model StarCoder2-3B --device cuda
Restart=always
RestartSec=10
Environment="TABBY_ROOT=/var/lib/tabby"
[Install]
WantedBy=multi-user.target
EOF
sudo systemctl enable tabby
sudo systemctl start tabby
Method 4: Docker Compose (Production)
# docker-compose.yml
version: '3.8'
services:
tabby:
image: tabbyml/tabby
command: serve --model StarCoder2-7B --device cuda
ports:
- "8080:8080"
volumes:
- tabby-data:/data
deploy:
resources:
reservations:
devices:
- driver: nvidia
count: 1
capabilities: [gpu]
restart: always
volumes:
tabby-data:
docker compose up -d
Model Selection Guide
Choosing the right model is the single most important decision for your Tabby setup. The tradeoff is always latency vs quality.
Available Models
At Q8 a model's weights occupy roughly 1 GB per billion parameters, plus a little for the KV cache and runtime, which is where the VRAM column below comes from. The same figure sets speed: generating a token means reading every weight out of VRAM, so on a given GPU, decode cost scales with parameter count. A 7B is about twice as slow per token as a 3B on the same card — that is arithmetic, not a benchmark, and it is the entire reason small models dominate autocomplete.
| Model | Parameters | VRAM (Q8) | Relative decode cost | Quality tier |
|---|---|---|---|---|
| StarCoder2-3B | 3B | 3.5 GB | 1x baseline | Good |
| StarCoder2-7B | 7B | 7.5 GB | ~2.3x | Very Good |
| StarCoder2-15B | 15B | 16 GB | ~5x | Excellent |
| DeepSeek-Coder 1.3B | 1.3B | 1.5 GB | ~0.4x | Basic |
| DeepSeek-Coder 6.7B | 6.7B | 7 GB | ~2.2x | Very Good |
| CodeLlama-7B | 7B | 7.5 GB | ~2.3x | Good |
| CodeLlama-13B | 13B | 14 GB | ~4.3x | Very Good |
| Qwen2.5-Coder-3B | 3B | 3.5 GB | 1x baseline | Good |
| Qwen2.5-Coder-7B | 7B | 7.5 GB | ~2.3x | Very Good |
Relative decode cost is parameter count divided by 3B, so it tells you how a swap will move your latency on hardware you already have. Quality tiers are editorial judgements based on each model's published evaluations and the size class it competes in, not scores we produced.
Which Model to Pick
Under 200ms is the target to design toward. Above roughly that, a developer notices the delay and starts typing ahead of the suggestion, which destroys the value of autocomplete regardless of how good the completion would have been. Since latency scales with model size, that target is really a VRAM-and-bandwidth budget:
- 4 GB VRAM (GTX 1070, RX 580): StarCoder2-3B or DeepSeek-Coder 1.3B
- 8 GB VRAM (RTX 3060 8GB, RTX 4060): StarCoder2-3B (fast) or DeepSeek-Coder 6.7B (quality)
- 12-16 GB VRAM (RTX 3060 12GB, RTX 4060 Ti 16GB): StarCoder2-7B (recommended sweet spot)
- 24 GB VRAM (RTX 4090): StarCoder2-7B with room for team serving, or StarCoder2-15B for single user
- Apple Silicon 16 GB: StarCoder2-3B (fast) or Qwen2.5-Coder-3B
- Apple Silicon 32 GB+: StarCoder2-7B or DeepSeek-Coder 6.7B
The sensible default for most teams: StarCoder2-7B on an RTX 4090. The 7B size is the last step up in quality you can take before decode cost roughly doubles again, and the 4090's 24 GB leaves the model (7.5 GB) plenty of headroom for per-request KV cache across several simultaneous developers — capacity on a completion server is bounded by VRAM headroom and bandwidth, not by a seat count.
Switching Models
# Stop current instance, start with new model
tabby serve --model DeepSeek-Coder-6.7B --device cuda
# Or in Docker
docker run -it --gpus all \
-p 8080:8080 \
-v $HOME/.tabby:/data \
tabbyml/tabby \
serve --model DeepSeek-Coder-6.7B --device cuda
Models are downloaded automatically on first use. A 7B model downloads ~7 GB on the first run.
Run this on your own machine and stop paying every month
Pay once and keep it. No renewal, no per-token bill, and nothing you feed it ever leaves your hardware.
IDE Integration
VS Code
# Install the extension
code --install-extension TabbyML.vscode-tabby
Configure in VS Code settings (Cmd+Shift+P > Preferences: Open Settings JSON):
{
"tabby.api.endpoint": "http://localhost:8080",
"tabby.api.token": "your-auth-token",
"tabby.inlineCompletion.triggerMode": "automatic",
"tabby.inlineCompletion.debounce": 200
}
Completions appear inline as you type, identical to Copilot. Press Tab to accept, Escape to dismiss.
JetBrains (IntelliJ, PyCharm, WebStorm, etc.)
- Open Settings > Plugins > Marketplace
- Search "Tabby" and install
- Settings > Tools > Tabby > Server Endpoint:
http://localhost:8080 - Enter your auth token
- Restart IDE
Vim / Neovim
" Using vim-plug
Plug 'TabbyML/vim-tabby'
" Configuration in .vimrc or init.vim
let g:tabby_server_url = 'http://localhost:8080'
let g:tabby_token = 'your-auth-token'
For Neovim with Lua config:
-- In init.lua
require('tabby').setup({
server_url = 'http://localhost:8080',
token = 'your-auth-token',
})
Verifying IDE Connection
After configuring any IDE, type some code and wait 200-300ms. If completions appear grayed out inline, the connection works. If not:
# Check server is running
curl http://localhost:8080/v1/health
# Test completion endpoint directly
curl -X POST http://localhost:8080/v1/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer your-auth-token" \
-d '{
"language": "python",
"segments": {
"prefix": "def fibonacci(n):\n if n <= 1:\n return n\n ",
"suffix": ""
}
}'
GPU Requirements and Performance
What Sets the Speed Ceiling
You can work out roughly what a given GPU is capable of without benchmarking anything, because token generation is memory-bandwidth bound: every token requires reading the whole weight set out of VRAM. At Q8 that is about 1 GB per billion parameters, so the ceiling is simply the card's published bandwidth divided by the model's size.
| GPU | Memory bandwidth | 3B at Q8 (~3 GB/token) | 7B at Q8 (~7 GB/token) |
|---|---|---|---|
| Apple M3 Pro (18 GB) | 150 GB/s | ~50 tok/s | ~21 tok/s |
| RTX 4060 | 272 GB/s | ~91 tok/s | ~39 tok/s |
| RTX 3060 12GB | 360 GB/s | ~120 tok/s | ~51 tok/s |
| RTX 4070 | 504 GB/s | ~168 tok/s | ~72 tok/s |
| RTX 3090 | 936 GB/s | ~312 tok/s | ~134 tok/s |
| RTX 4090 | 1,008 GB/s | ~336 tok/s | ~144 tok/s |
| RTX 5090 | 1,792 GB/s | ~597 tok/s | ~256 tok/s |
These are arithmetic upper bounds, not measurements. Any real deployment lands below them, because the numbers ignore prefill (reading the surrounding code into the context), attention over the KV cache, and framework overhead. What they are good for is comparing cards and sizing a purchase: the ratios hold even though the absolute values do not.
Turning tok/s into completions: a completion is not one token. An inline suggestion is typically one line to a short block, so call it a few dozen tokens. Divide the ceiling above by that, subtract prefill time, then add your IDE's debounce (200ms by default) and network round trip — and you have a rough envelope for what a single developer will feel and how many developers a server can carry. Once the server is live, stop estimating: the admin dashboard reports real latency percentiles for your actual model, hardware, and codebase.
Power Consumption
Vendor TDP is the honest input here, since it is the number NVIDIA and Apple publish and the ceiling your card can draw: RTX 3060 12GB is 170W, RTX 4090 is 450W, RTX 5090 is 575W, and an A6000 is 300W. Apple does not publish a comparable GPU TDP for Apple Silicon, but the whole-package draw of a Mac laptop or Mini is a fraction of any of those, which is the real reason Apple Silicon is attractive for an always-on completion server.
Cost follows directly: watts x hours ÷ 1000 x your rate per kWh. An RTX 4090 pinned at its full 450W through a 176-hour work month is 79 kWh, or about $9.50 at $0.12/kWh — and that is a hard ceiling, not a forecast. An autocomplete server spends most of its day idle between keystrokes, so the real bill is a fraction of it. Run the same arithmetic with your own electricity rate before it factors into a buying decision.
Repository Indexing
One of Tabby's strongest features: it can index your private repositories and use that context to improve completions. Instead of generic code suggestions, you get completions that match your project's patterns, API usage, and naming conventions.
Setting Up Repository Indexing
- Open the Tabby admin panel at
http://localhost:8080 - Navigate to Settings > Repositories
- Add your Git repository:
# Through the admin API
curl -X POST http://localhost:8080/v1/repositories \
-H "Content-Type: application/json" \
-H "Authorization: Bearer admin-token" \
-d '{
"name": "my-project",
"git_url": "file:///path/to/my-project",
"branch": "main"
}'
- Tabby clones the repository and builds a code index in the background
- Once indexed, completions automatically incorporate your codebase patterns
What Indexing Improves
- Import suggestions: Tabby learns which modules your project uses and suggests correct imports
- Function signatures: Completions match your naming conventions (camelCase, snake_case, etc.)
- API patterns: If your codebase always calls
db.query().where().first(), Tabby suggests that chain - Type patterns: Consistent with your TypeScript types, Python type hints, etc.
What Indexing Costs You
Indexing is a one-time pass over the repository plus incremental updates afterwards, so build time scales with the amount of code, and the resulting index scales with it too. There is no universal number to quote — it depends on your language mix, file sizes, and disk speed as much as on line count — but three things are reliably true and worth planning around:
- The build is background work. Completions keep serving while the index is being built, so you are not choosing between the two.
- The index lives in memory alongside the model. That overhead is additive to model VRAM, so a card that fits your model exactly is a card with no room for indexing.
- Only the first pass is expensive. After that Tabby updates incrementally, so the cost you pay once is not the cost you pay daily.
Index one repository first and watch what it actually costs on your hardware before pointing Tabby at a monorepo.
Tabby vs Continue.dev
Both Tabby and Continue.dev are open-source tools for local AI coding, but they serve different purposes. For a detailed Continue.dev setup, see our Continue.dev + Ollama guide.
| Feature | Tabby | Continue.dev + Ollama |
|---|---|---|
| Primary purpose | Code completion (autocomplete) | Full AI coding assistant |
| Tab autocomplete | Excellent (purpose-built) | Good (secondary feature) |
| Chat | No | Yes |
| Edit mode | No | Yes |
| Agent mode | No | Yes |
| Model hosting | Built-in | Requires Ollama |
| Repository indexing | Built-in | Via embeddings model |
| Team features | Built-in (auth, analytics) | None |
| IDE support | VS Code, JetBrains, Vim | VS Code, JetBrains |
| Setup complexity | One command | Ollama + Continue + config |
| Resource usage | Low (one model) | Higher (autocomplete + chat models) |
The Ideal Setup: Use Both
The best local AI coding setup combines them:
- Tabby for real-time autocomplete (StarCoder2-3B, ~3.5 GB VRAM)
- Continue.dev with Ollama for chat, debugging, and refactoring (Qwen2.5-Coder 7B, ~7.5 GB VRAM)
- Total VRAM: ~11 GB, fits on RTX 3060 12GB or RTX 4060 Ti 16GB
That covers both halves of the job — inline autocomplete and a chat/edit assistant — with nothing leaving your machine.
Configure Continue.dev to not use its own autocomplete (since Tabby handles that):
# ~/.continue/config.yaml
models:
- name: Qwen2.5-Coder 7B
provider: ollama
model: qwen2.5-coder:7b
roles:
- chat
- edit
- apply
# No autocomplete model - Tabby handles it
Team Deployment
Authentication Setup
By default, Tabby runs without authentication. For team deployments, enable auth:
# Start with authentication enabled
tabby serve --model StarCoder2-7B --device cuda
# On first run, create admin account at http://your-server:8080
# Then invite team members through the admin panel
Tabby supports:
- Built-in email/password authentication
- OAuth (GitHub, Google, GitLab)
- LDAP (enterprise)
Network Configuration
For team access, bind Tabby to your LAN:
# Bind to all interfaces
tabby serve --model StarCoder2-7B --device cuda --host 0.0.0.0 --port 8080
# Or use a reverse proxy (nginx)
# /etc/nginx/sites-available/tabby
server {
listen 443 ssl;
server_name tabby.internal.company.com;
ssl_certificate /etc/ssl/certs/internal.crt;
ssl_certificate_key /etc/ssl/private/internal.key;
location / {
proxy_pass http://localhost:8080;
proxy_set_header Host $host;
proxy_set_header X-Real-IP $remote_addr;
}
}
Monitoring and Analytics
The Tabby admin dashboard (http://your-server:8080) shows:
- Active users and sessions
- Completion acceptance rate (how often developers press Tab)
- Completions per hour/day
- Model latency percentiles
- GPU utilization
Two of those are worth watching closely:
- Acceptance rate is the honest quality signal, because it is developers voting with the Tab key. The absolute number matters far less than the direction: record your baseline, change one thing (a bigger model, repository indexing, a different debounce), and see whether it moves. That is a real experiment on your own codebase, which beats any published figure.
- Latency percentiles are what to tune against. Watch P95 rather than the average — the average hides the slow requests, and the slow ones are the ones a developer types straight past. If P95 drifts above the ~200ms threshold where suggestions stop being useful, drop to a smaller model or raise the debounce.
GPU utilization is the capacity gauge. If it sits high through the working day, you are near the ceiling and it is time to add a card or a smaller model.
Scaling for Larger Teams
For 50+ developers, consider:
- Horizontal scaling: Run multiple Tabby instances behind a load balancer
# Instance 1 on GPU 0
CUDA_VISIBLE_DEVICES=0 tabby serve --model StarCoder2-7B --port 8080
# Instance 2 on GPU 1
CUDA_VISIBLE_DEVICES=1 tabby serve --model StarCoder2-7B --port 8081
-
Dedicated hardware: An NVIDIA A6000 48GB or dual RTX 4090s buy you both more bandwidth and far more VRAM headroom for concurrent KV caches, which are the two things that actually cap a completion server
-
Cloud deployment: Deploy on a cloud GPU instance (RunPod, Lambda) if you do not want on-premise hardware
Performance Tuning
Reduce Latency
# Use a smaller model for faster completions
tabby serve --model StarCoder2-3B --device cuda
# Limit completion length (fewer tokens = faster)
# In the admin panel: Settings > Completion > Max tokens: 128
# Increase GPU memory allocation
CUDA_MEM_FRACTION=0.9 tabby serve --model StarCoder2-7B --device cuda
IDE-Side Tuning
In VS Code settings:
{
"tabby.inlineCompletion.debounce": 250,
"tabby.inlineCompletion.triggerMode": "automatic",
"tabby.maxPrefixLines": 20,
"tabby.maxSuffixLines": 20
}
Increasing debounce from 200ms to 250-300ms reduces server load (fewer requests) at the cost of slightly delayed suggestions. For slow servers, this trade is worth it.
Model Warm-Up
The first completion after a cold start is slow because the model loads into GPU memory. Keep the model warm:
# Send periodic health checks to prevent unloading
while true; do
curl -s http://localhost:8080/v1/health > /dev/null
sleep 60
done &
Monitoring GPU Usage
# Watch GPU utilization in real-time
watch -n 1 nvidia-smi
# Or use nvtop for a better visualization
nvtop
If GPU utilization is consistently above 80%, either add another GPU, switch to a smaller model, or increase the debounce time on IDE clients.
Privacy Advantages
The core value proposition of Tabby over cloud services deserves emphasis:
-
Zero data exfiltration: Your code stays on your hardware. Period. No telemetry, no training data collection, no third-party access.
-
Compliance friendly: Self-hosted satisfies SOC 2, HIPAA, GDPR, and FedRAMP data residency requirements. Your security team will appreciate not having to review another SaaS vendor's DPA.
-
No vendor lock-in: Switch models, modify the source code, or migrate to different hardware anytime. Apache 2.0 license means you own the deployment.
-
Air-gapped support: Tabby runs entirely offline after the initial model download. Disconnect from the internet and it works identically. Critical for defense, government, and high-security environments.
For teams already running local AI for other tasks, see our local AI programming models guide for complementary tools.
Conclusion
Tabby is the most mature self-hosted code completion server available, and its scope is the reason: it does autocomplete, it does it from a model you control, and it leaves everything else to other tools. The privacy guarantee is the part that is genuinely absolute — the code never leaves the hardware you own. Completion quality is not absolute, and it is not really Tabby's to claim: it belongs to the 3B-7B model you load, so it improves when you spend VRAM, not when you upgrade the server.
For a single developer, the ROI calculation depends on how much you value privacy. For a team of 10+, the math is clear: one RTX 4090 ($1,600) replaces $1,200/year in Copilot subscriptions and eliminates code leaving your network.
Start with Docker and StarCoder2-3B, and judge it on your acceptance rate rather than on anyone's benchmark. If the completions feel useful, upgrade to StarCoder2-7B. Index your repositories for the biggest quality improvement. And combine it with Continue.dev + Ollama for a complete local AI coding stack that matches what the cloud providers charge $20-50/month per seat to deliver.
Building a complete local AI development environment? Check our best local AI coding models ranking or the AI hardware requirements guide to plan your hardware.
Picked your coding model? Build a real AI dev workflow.
From local copilots to agents that ship code — the structured path, running on your hardware. First chapter free.
Liked this? 25 full AI courses are waiting.
From fundamentals to RAG, agents, MCP servers, voice AI, and production deployment with real GitHub repos. First chapter free, every course.
Build Real AI on Your Machine
RAG, agents, NLP, vision, and MLOps - chapters across 25 courses that take you from reading about AI to building AI.
Want the structured version?
Hands-on courses on local AI, from $8.99 a month. The first chapter of each is free.
Keep going
- PILLARBest Local AI for Coding 2026: 10 Models Ranked by VRAM
- 7B vs 14B vs 32B vs 70B for Coding (2026): What Size?
- AI Context Windows: 4K vs 128K vs 1M Tokens Explained (2026)
- Aider + Ollama Setup (2026): Free Local AI Coding Agent
- Best 14B Coding Models (2026): Ranked by HumanEval + VRAM
- Best AI Coding Models Ranked: SWE-bench Leaderboard
- Best AI for JavaScript & TypeScript 2026: 10 Models Ranked
- Best AI Models for Python Development 2026: Top 10 Ranked
- Best Claude Model for Coding: Sonnet 5, Opus 4.8, Fable 5
- Best Ollama Model for Coding (2026): Qwen3-Coder Ranked #1
Comments (0)
No comments yet. Be the first to share your thoughts!