★ Reading this for free? Get 25 structured AI courses + per-chapter AI tutor — the first chapter of every course free, no card.Start free in 30 secondsOr own every course: $149 once
Tools

Tabby: Self-Hosted GitHub Copilot Alternative

April 10, 2026
19 min read
Local AI Master Research Team

Want to go deeper than this article?

Free account unlocks the first chapter of all 25 courses — RAG, agents, MCP, voice AI, MLOps, real GitHub repos.

📚AI Learning Path

Picked your coding model? Build a real AI dev workflow. From local copilots to agents that ship code — the structured path, running on your hardware. First chapter free.

Start free
Or own it for life — Lifetime $149, pay once

Published on April 10, 2026 -- 19 min read

GitHub Copilot costs $10/month per developer and sends every keystroke to Microsoft's servers. For a 10-person team, that is $1,200/year -- and your proprietary code passes through infrastructure you do not control. Tabby eliminates both problems: it is free, open-source, and runs entirely on your network.

Tabby has 23,000+ GitHub stars, supports VS Code, JetBrains, Vim, and Neovim, and serves real-time code completions from models you choose. The architecture is deliberately small: one server process holding one model in GPU memory, an HTTP endpoint, and an editor extension that calls it. That is why the install is short and the failure modes are few.

Set your expectations by the model, not by the server. Tabby is plumbing — the suggestions you get are whatever a 3B to 7B open code model produces. Those models are strong at the completions that make up most of autocomplete's value (finishing a line, filling a known pattern, matching the shape of the code around the cursor) and weaker at long, novel blocks, which is where a frontier cloud model still pulls ahead.

Here is how to set it up, which model to pick, and how to tune it for your team.


What is Tabby

Tabby is an open-source AI code completion server built by TabbyML. It provides:

  • Real-time inline completion served over HTTP from a model held in local GPU memory
  • Multi-IDE support: VS Code, JetBrains (IntelliJ, PyCharm, WebStorm), Vim, Neovim
  • Model flexibility: StarCoder2, DeepSeek-Coder, CodeLlama, Qwen2.5-Coder
  • Repository indexing: learns your codebase patterns for better suggestions
  • Admin dashboard: user management, usage analytics, model configuration
  • Enterprise features: LDAP/OAuth auth, audit logging, access controls

The architecture is simple: Tabby runs as a server (standalone binary, Docker container, or Homebrew install), loads a code model into GPU memory, and serves completions over HTTP. IDE extensions connect to the server and inject suggestions inline, identical to how Copilot works.

What Tabby Does Not Do

Tabby focuses specifically on code completion -- the autocomplete experience. It does not include:

  • Chat interface (use Continue.dev or Claude for that)
  • Code explanation or documentation generation
  • Agent mode or autonomous task execution
  • Code review or PR analysis

This focused scope is actually an advantage: Tabby does one thing and does it well, with minimal resource usage.


Reading articles is good. Building is better.

Free account = the first chapter of all 25 courses, with a per-chapter AI tutor. No card.

Why Self-Host Code Completion

Privacy

Every character you type in Copilot gets sent to GitHub's servers. For companies handling regulated data (healthcare, finance, defense), customer PII, or proprietary algorithms, that is a compliance problem. Self-hosted Tabby keeps all code on your network.

Cost at Scale

Team SizeCopilot Cost/YearTabby Hardware CostTabby Breakeven
5 devs$600$1,600 (RTX 4090)2.7 years
10 devs$1,200$1,600 (RTX 4090)1.3 years
25 devs$3,000$1,600 (RTX 4090)6.4 months
50 devs$6,000$4,500 (A6000 48GB)9 months
100 devs$12,000$4,500 (A6000 48GB)4.5 months

At 10+ developers, Tabby pays for itself in the first year. At 25+, the savings are substantial.

Customization

Copilot gives you one model, take it or leave it. Tabby lets you:

  • Choose models optimized for your languages (DeepSeek-Coder for Python, StarCoder2 for polyglot)
  • Index your private repositories for context-aware completions
  • Fine-tune models on your codebase (advanced)
  • Control context window size and completion behavior

Uptime Independence

Copilot goes down when GitHub has an outage. Your Tabby server runs on your infrastructure, on your schedule.


Installation Methods

Docker is the cleanest way to deploy Tabby, especially for team servers.

# NVIDIA GPU (CUDA)
docker run -it \
  --gpus all \
  -p 8080:8080 \
  -v $HOME/.tabby:/data \
  tabbyml/tabby \
  serve --model StarCoder2-3B --device cuda

# AMD GPU (ROCm)
docker run -it \
  --device /dev/kfd --device /dev/dri \
  --group-add video \
  -p 8080:8080 \
  -v $HOME/.tabby:/data \
  tabbyml/tabby-rocm \
  serve --model StarCoder2-3B --device rocm

After startup, open http://localhost:8080 in your browser. You will see the admin dashboard where you can create user accounts, manage models, and view usage analytics.

Method 2: Homebrew (macOS)

# Install
brew install tabbyml/tabby/tabby

# Run with Apple Metal acceleration
tabby serve --model StarCoder2-3B --device metal

# Verify it is running
curl http://localhost:8080/v1/health

Method 3: Direct Binary (Linux)

# Download the latest release
curl -L https://github.com/TabbyML/tabby/releases/latest/download/tabby_x86_64-unknown-linux-gnu -o tabby
chmod +x tabby

# Run with CUDA
./tabby serve --model StarCoder2-3B --device cuda

# Or run as a systemd service for persistence
sudo tee /etc/systemd/system/tabby.service << 'EOF'
[Unit]
Description=Tabby AI Code Completion Server
After=network.target

[Service]
Type=simple
User=tabby
ExecStart=/usr/local/bin/tabby serve --model StarCoder2-3B --device cuda
Restart=always
RestartSec=10
Environment="TABBY_ROOT=/var/lib/tabby"

[Install]
WantedBy=multi-user.target
EOF

sudo systemctl enable tabby
sudo systemctl start tabby

Method 4: Docker Compose (Production)

# docker-compose.yml
version: '3.8'
services:
  tabby:
    image: tabbyml/tabby
    command: serve --model StarCoder2-7B --device cuda
    ports:
      - "8080:8080"
    volumes:
      - tabby-data:/data
    deploy:
      resources:
        reservations:
          devices:
            - driver: nvidia
              count: 1
              capabilities: [gpu]
    restart: always

volumes:
  tabby-data:
docker compose up -d

Model Selection Guide

Choosing the right model is the single most important decision for your Tabby setup. The tradeoff is always latency vs quality.

Available Models

At Q8 a model's weights occupy roughly 1 GB per billion parameters, plus a little for the KV cache and runtime, which is where the VRAM column below comes from. The same figure sets speed: generating a token means reading every weight out of VRAM, so on a given GPU, decode cost scales with parameter count. A 7B is about twice as slow per token as a 3B on the same card — that is arithmetic, not a benchmark, and it is the entire reason small models dominate autocomplete.

ModelParametersVRAM (Q8)Relative decode costQuality tier
StarCoder2-3B3B3.5 GB1x baselineGood
StarCoder2-7B7B7.5 GB~2.3xVery Good
StarCoder2-15B15B16 GB~5xExcellent
DeepSeek-Coder 1.3B1.3B1.5 GB~0.4xBasic
DeepSeek-Coder 6.7B6.7B7 GB~2.2xVery Good
CodeLlama-7B7B7.5 GB~2.3xGood
CodeLlama-13B13B14 GB~4.3xVery Good
Qwen2.5-Coder-3B3B3.5 GB1x baselineGood
Qwen2.5-Coder-7B7B7.5 GB~2.3xVery Good

Relative decode cost is parameter count divided by 3B, so it tells you how a swap will move your latency on hardware you already have. Quality tiers are editorial judgements based on each model's published evaluations and the size class it competes in, not scores we produced.

Which Model to Pick

Under 200ms is the target to design toward. Above roughly that, a developer notices the delay and starts typing ahead of the suggestion, which destroys the value of autocomplete regardless of how good the completion would have been. Since latency scales with model size, that target is really a VRAM-and-bandwidth budget:

  • 4 GB VRAM (GTX 1070, RX 580): StarCoder2-3B or DeepSeek-Coder 1.3B
  • 8 GB VRAM (RTX 3060 8GB, RTX 4060): StarCoder2-3B (fast) or DeepSeek-Coder 6.7B (quality)
  • 12-16 GB VRAM (RTX 3060 12GB, RTX 4060 Ti 16GB): StarCoder2-7B (recommended sweet spot)
  • 24 GB VRAM (RTX 4090): StarCoder2-7B with room for team serving, or StarCoder2-15B for single user
  • Apple Silicon 16 GB: StarCoder2-3B (fast) or Qwen2.5-Coder-3B
  • Apple Silicon 32 GB+: StarCoder2-7B or DeepSeek-Coder 6.7B

The sensible default for most teams: StarCoder2-7B on an RTX 4090. The 7B size is the last step up in quality you can take before decode cost roughly doubles again, and the 4090's 24 GB leaves the model (7.5 GB) plenty of headroom for per-request KV cache across several simultaneous developers — capacity on a completion server is bounded by VRAM headroom and bandwidth, not by a seat count.

Switching Models

# Stop current instance, start with new model
tabby serve --model DeepSeek-Coder-6.7B --device cuda

# Or in Docker
docker run -it --gpus all \
  -p 8080:8080 \
  -v $HOME/.tabby:/data \
  tabbyml/tabby \
  serve --model DeepSeek-Coder-6.7B --device cuda

Models are downloaded automatically on first use. A 7B model downloads ~7 GB on the first run.


Own it instead of renting it

Run this on your own machine and stop paying every month

Pay once and keep it. No renewal, no per-token bill, and nothing you feed it ever leaves your hardware.

IDE Integration

VS Code

# Install the extension
code --install-extension TabbyML.vscode-tabby

Configure in VS Code settings (Cmd+Shift+P > Preferences: Open Settings JSON):

{
  "tabby.api.endpoint": "http://localhost:8080",
  "tabby.api.token": "your-auth-token",
  "tabby.inlineCompletion.triggerMode": "automatic",
  "tabby.inlineCompletion.debounce": 200
}

Completions appear inline as you type, identical to Copilot. Press Tab to accept, Escape to dismiss.

JetBrains (IntelliJ, PyCharm, WebStorm, etc.)

  1. Open Settings > Plugins > Marketplace
  2. Search "Tabby" and install
  3. Settings > Tools > Tabby > Server Endpoint: http://localhost:8080
  4. Enter your auth token
  5. Restart IDE

Vim / Neovim

" Using vim-plug
Plug 'TabbyML/vim-tabby'

" Configuration in .vimrc or init.vim
let g:tabby_server_url = 'http://localhost:8080'
let g:tabby_token = 'your-auth-token'

For Neovim with Lua config:

-- In init.lua
require('tabby').setup({
  server_url = 'http://localhost:8080',
  token = 'your-auth-token',
})

Verifying IDE Connection

After configuring any IDE, type some code and wait 200-300ms. If completions appear grayed out inline, the connection works. If not:

# Check server is running
curl http://localhost:8080/v1/health

# Test completion endpoint directly
curl -X POST http://localhost:8080/v1/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer your-auth-token" \
  -d '{
    "language": "python",
    "segments": {
      "prefix": "def fibonacci(n):\n    if n <= 1:\n        return n\n    ",
      "suffix": ""
    }
  }'

GPU Requirements and Performance

What Sets the Speed Ceiling

You can work out roughly what a given GPU is capable of without benchmarking anything, because token generation is memory-bandwidth bound: every token requires reading the whole weight set out of VRAM. At Q8 that is about 1 GB per billion parameters, so the ceiling is simply the card's published bandwidth divided by the model's size.

GPUMemory bandwidth3B at Q8 (~3 GB/token)7B at Q8 (~7 GB/token)
Apple M3 Pro (18 GB)150 GB/s~50 tok/s~21 tok/s
RTX 4060272 GB/s~91 tok/s~39 tok/s
RTX 3060 12GB360 GB/s~120 tok/s~51 tok/s
RTX 4070504 GB/s~168 tok/s~72 tok/s
RTX 3090936 GB/s~312 tok/s~134 tok/s
RTX 40901,008 GB/s~336 tok/s~144 tok/s
RTX 50901,792 GB/s~597 tok/s~256 tok/s

These are arithmetic upper bounds, not measurements. Any real deployment lands below them, because the numbers ignore prefill (reading the surrounding code into the context), attention over the KV cache, and framework overhead. What they are good for is comparing cards and sizing a purchase: the ratios hold even though the absolute values do not.

Turning tok/s into completions: a completion is not one token. An inline suggestion is typically one line to a short block, so call it a few dozen tokens. Divide the ceiling above by that, subtract prefill time, then add your IDE's debounce (200ms by default) and network round trip — and you have a rough envelope for what a single developer will feel and how many developers a server can carry. Once the server is live, stop estimating: the admin dashboard reports real latency percentiles for your actual model, hardware, and codebase.

Power Consumption

Vendor TDP is the honest input here, since it is the number NVIDIA and Apple publish and the ceiling your card can draw: RTX 3060 12GB is 170W, RTX 4090 is 450W, RTX 5090 is 575W, and an A6000 is 300W. Apple does not publish a comparable GPU TDP for Apple Silicon, but the whole-package draw of a Mac laptop or Mini is a fraction of any of those, which is the real reason Apple Silicon is attractive for an always-on completion server.

Cost follows directly: watts x hours ÷ 1000 x your rate per kWh. An RTX 4090 pinned at its full 450W through a 176-hour work month is 79 kWh, or about $9.50 at $0.12/kWh — and that is a hard ceiling, not a forecast. An autocomplete server spends most of its day idle between keystrokes, so the real bill is a fraction of it. Run the same arithmetic with your own electricity rate before it factors into a buying decision.


Repository Indexing

One of Tabby's strongest features: it can index your private repositories and use that context to improve completions. Instead of generic code suggestions, you get completions that match your project's patterns, API usage, and naming conventions.

Setting Up Repository Indexing

  1. Open the Tabby admin panel at http://localhost:8080
  2. Navigate to Settings > Repositories
  3. Add your Git repository:
# Through the admin API
curl -X POST http://localhost:8080/v1/repositories \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer admin-token" \
  -d '{
    "name": "my-project",
    "git_url": "file:///path/to/my-project",
    "branch": "main"
  }'
  1. Tabby clones the repository and builds a code index in the background
  2. Once indexed, completions automatically incorporate your codebase patterns

What Indexing Improves

  • Import suggestions: Tabby learns which modules your project uses and suggests correct imports
  • Function signatures: Completions match your naming conventions (camelCase, snake_case, etc.)
  • API patterns: If your codebase always calls db.query().where().first(), Tabby suggests that chain
  • Type patterns: Consistent with your TypeScript types, Python type hints, etc.

What Indexing Costs You

Indexing is a one-time pass over the repository plus incremental updates afterwards, so build time scales with the amount of code, and the resulting index scales with it too. There is no universal number to quote — it depends on your language mix, file sizes, and disk speed as much as on line count — but three things are reliably true and worth planning around:

  • The build is background work. Completions keep serving while the index is being built, so you are not choosing between the two.
  • The index lives in memory alongside the model. That overhead is additive to model VRAM, so a card that fits your model exactly is a card with no room for indexing.
  • Only the first pass is expensive. After that Tabby updates incrementally, so the cost you pay once is not the cost you pay daily.

Index one repository first and watch what it actually costs on your hardware before pointing Tabby at a monorepo.


Tabby vs Continue.dev

Both Tabby and Continue.dev are open-source tools for local AI coding, but they serve different purposes. For a detailed Continue.dev setup, see our Continue.dev + Ollama guide.

FeatureTabbyContinue.dev + Ollama
Primary purposeCode completion (autocomplete)Full AI coding assistant
Tab autocompleteExcellent (purpose-built)Good (secondary feature)
ChatNoYes
Edit modeNoYes
Agent modeNoYes
Model hostingBuilt-inRequires Ollama
Repository indexingBuilt-inVia embeddings model
Team featuresBuilt-in (auth, analytics)None
IDE supportVS Code, JetBrains, VimVS Code, JetBrains
Setup complexityOne commandOllama + Continue + config
Resource usageLow (one model)Higher (autocomplete + chat models)

The Ideal Setup: Use Both

The best local AI coding setup combines them:

  1. Tabby for real-time autocomplete (StarCoder2-3B, ~3.5 GB VRAM)
  2. Continue.dev with Ollama for chat, debugging, and refactoring (Qwen2.5-Coder 7B, ~7.5 GB VRAM)
  3. Total VRAM: ~11 GB, fits on RTX 3060 12GB or RTX 4060 Ti 16GB

That covers both halves of the job — inline autocomplete and a chat/edit assistant — with nothing leaving your machine.

Configure Continue.dev to not use its own autocomplete (since Tabby handles that):

# ~/.continue/config.yaml
models:
  - name: Qwen2.5-Coder 7B
    provider: ollama
    model: qwen2.5-coder:7b
    roles:
      - chat
      - edit
      - apply
# No autocomplete model - Tabby handles it

Team Deployment

Authentication Setup

By default, Tabby runs without authentication. For team deployments, enable auth:

# Start with authentication enabled
tabby serve --model StarCoder2-7B --device cuda

# On first run, create admin account at http://your-server:8080
# Then invite team members through the admin panel

Tabby supports:

  • Built-in email/password authentication
  • OAuth (GitHub, Google, GitLab)
  • LDAP (enterprise)

Network Configuration

For team access, bind Tabby to your LAN:

# Bind to all interfaces
tabby serve --model StarCoder2-7B --device cuda --host 0.0.0.0 --port 8080

# Or use a reverse proxy (nginx)
# /etc/nginx/sites-available/tabby
server {
    listen 443 ssl;
    server_name tabby.internal.company.com;

    ssl_certificate /etc/ssl/certs/internal.crt;
    ssl_certificate_key /etc/ssl/private/internal.key;

    location / {
        proxy_pass http://localhost:8080;
        proxy_set_header Host $host;
        proxy_set_header X-Real-IP $remote_addr;
    }
}

Monitoring and Analytics

The Tabby admin dashboard (http://your-server:8080) shows:

  • Active users and sessions
  • Completion acceptance rate (how often developers press Tab)
  • Completions per hour/day
  • Model latency percentiles
  • GPU utilization

Two of those are worth watching closely:

  • Acceptance rate is the honest quality signal, because it is developers voting with the Tab key. The absolute number matters far less than the direction: record your baseline, change one thing (a bigger model, repository indexing, a different debounce), and see whether it moves. That is a real experiment on your own codebase, which beats any published figure.
  • Latency percentiles are what to tune against. Watch P95 rather than the average — the average hides the slow requests, and the slow ones are the ones a developer types straight past. If P95 drifts above the ~200ms threshold where suggestions stop being useful, drop to a smaller model or raise the debounce.

GPU utilization is the capacity gauge. If it sits high through the working day, you are near the ceiling and it is time to add a card or a smaller model.

Scaling for Larger Teams

For 50+ developers, consider:

  1. Horizontal scaling: Run multiple Tabby instances behind a load balancer
# Instance 1 on GPU 0
CUDA_VISIBLE_DEVICES=0 tabby serve --model StarCoder2-7B --port 8080

# Instance 2 on GPU 1
CUDA_VISIBLE_DEVICES=1 tabby serve --model StarCoder2-7B --port 8081
  1. Dedicated hardware: An NVIDIA A6000 48GB or dual RTX 4090s buy you both more bandwidth and far more VRAM headroom for concurrent KV caches, which are the two things that actually cap a completion server

  2. Cloud deployment: Deploy on a cloud GPU instance (RunPod, Lambda) if you do not want on-premise hardware


Performance Tuning

Reduce Latency

# Use a smaller model for faster completions
tabby serve --model StarCoder2-3B --device cuda

# Limit completion length (fewer tokens = faster)
# In the admin panel: Settings > Completion > Max tokens: 128

# Increase GPU memory allocation
CUDA_MEM_FRACTION=0.9 tabby serve --model StarCoder2-7B --device cuda

IDE-Side Tuning

In VS Code settings:

{
  "tabby.inlineCompletion.debounce": 250,
  "tabby.inlineCompletion.triggerMode": "automatic",
  "tabby.maxPrefixLines": 20,
  "tabby.maxSuffixLines": 20
}

Increasing debounce from 200ms to 250-300ms reduces server load (fewer requests) at the cost of slightly delayed suggestions. For slow servers, this trade is worth it.

Model Warm-Up

The first completion after a cold start is slow because the model loads into GPU memory. Keep the model warm:

# Send periodic health checks to prevent unloading
while true; do
  curl -s http://localhost:8080/v1/health > /dev/null
  sleep 60
done &

Monitoring GPU Usage

# Watch GPU utilization in real-time
watch -n 1 nvidia-smi

# Or use nvtop for a better visualization
nvtop

If GPU utilization is consistently above 80%, either add another GPU, switch to a smaller model, or increase the debounce time on IDE clients.


Privacy Advantages

The core value proposition of Tabby over cloud services deserves emphasis:

  1. Zero data exfiltration: Your code stays on your hardware. Period. No telemetry, no training data collection, no third-party access.

  2. Compliance friendly: Self-hosted satisfies SOC 2, HIPAA, GDPR, and FedRAMP data residency requirements. Your security team will appreciate not having to review another SaaS vendor's DPA.

  3. No vendor lock-in: Switch models, modify the source code, or migrate to different hardware anytime. Apache 2.0 license means you own the deployment.

  4. Air-gapped support: Tabby runs entirely offline after the initial model download. Disconnect from the internet and it works identically. Critical for defense, government, and high-security environments.

For teams already running local AI for other tasks, see our local AI programming models guide for complementary tools.


Conclusion

Tabby is the most mature self-hosted code completion server available, and its scope is the reason: it does autocomplete, it does it from a model you control, and it leaves everything else to other tools. The privacy guarantee is the part that is genuinely absolute — the code never leaves the hardware you own. Completion quality is not absolute, and it is not really Tabby's to claim: it belongs to the 3B-7B model you load, so it improves when you spend VRAM, not when you upgrade the server.

For a single developer, the ROI calculation depends on how much you value privacy. For a team of 10+, the math is clear: one RTX 4090 ($1,600) replaces $1,200/year in Copilot subscriptions and eliminates code leaving your network.

Start with Docker and StarCoder2-3B, and judge it on your acceptance rate rather than on anyone's benchmark. If the completions feel useful, upgrade to StarCoder2-7B. Index your repositories for the biggest quality improvement. And combine it with Continue.dev + Ollama for a complete local AI coding stack that matches what the cloud providers charge $20-50/month per seat to deliver.


Building a complete local AI development environment? Check our best local AI coding models ranking or the AI hardware requirements guide to plan your hardware.

🎯
AI Learning Path

Picked your coding model? Build a real AI dev workflow.

From local copilots to agents that ship code — the structured path, running on your hardware. First chapter free.

Or own it for life — Lifetime $149 $599, pay once

Liked this? 25 full AI courses are waiting.

From fundamentals to RAG, agents, MCP servers, voice AI, and production deployment with real GitHub repos. First chapter free, every course.

Reading now
Join the discussion
TagsTabbySelf-HostedCode CompletionCopilot AlternativeOpen SourceVS CodeDocker

Local AI Master Research Team

Local AI Master writes hands-on courses and hardware guides for running AI on machines you own. Content is checked against current releases and corrected when readers tell us it is wrong.

Build Real AI on Your Machine

RAG, agents, NLP, vision, and MLOps - chapters across 25 courses that take you from reading about AI to building AI.

Want the structured version?

Hands-on courses on local AI, from $8.99 a month. The first chapter of each is free.

AI Learning Path
More on AI Models for Coding
See the full Best Local AI for Coding guide.

Comments (0)

No comments yet. Be the first to share your thoughts!

📅 Published: April 10, 2026🔄 Last Updated: April 10, 2026✓ Manually Reviewed
LM

Written by the Local AI Master Team

The team behind Local AI Master

We build Local AI Master around practical, testable local AI workflows: model selection, hardware planning, RAG systems, agents, and MLOps. The goal is to turn scattered tutorials into a structured learning path you can follow on your own hardware.

✓ Local AI Curriculum✓ Hands-On Projects✓ Open Source Contributor

Build Real AI on Your Machine

RAG, agents, NLP, vision, and MLOps - chapters across 25 courses that take you from reading about AI to building AI.

Was this helpful?

📚
Free · no account required

Grab the AI Starter Kit — career roadmap, cheat sheet, setup guide

No spam. Unsubscribe with one click.

🎯
AI Learning Path

Picked your coding model? Build a real AI dev workflow.

From local copilots to agents that ship code — the structured path, running on your hardware. First chapter free.

Or own it for life — Lifetime $149 $599, pay once
Free Tools & Calculators