---
title: "Installing the Ollama LLM with Docker and Benchmarking CPU vs GPU Hardware Acceleration"
description: "Deploy Ollama + Open WebUI on Windows WSL2 with Docker Compose, then benchmark CPU vs GPU (RTX 3080) inference speed and walk through configuring the NVIDIA Container Toolkit."
canonical_url: "https://blog.markkulab.net/en/post/build-your-ollama-ai-with-hardware-acceleration"
author: "Mark Ku"
author_url: "https://blog.markkulab.net/en/author/mark-ku"
site: "Mark Ku's Tech Notes"
date_published: "2024-08-27 01:01:35 +0800"
category: "AI"
tags: ["ollama", "ai", "gpu", "docker", "llm", "nvidia"]
language: "en"
license: "CC BY 4.0"
license_url: "https://creativecommons.org/licenses/by/4.0/"
attribution: "when reusing or quoting, credit the author and link back to the original"
---

# Installing the Ollama LLM with Docker and Benchmarking CPU vs GPU Hardware Acceleration

## Background
We previously used [Azure OpenAI to optimise SEO](https://blog.markkulab.net/build-your-ollama-ai-with-hardware-acceleration/), but Azure OpenAI is capped at $150 USD per month and we kept getting cut off near month-end when our MSDN credit ran out. So we started looking for an alternative and ended up choosing Ollama, the open-source project from Facebook.

## Installation and Setup
### First, create the Docker Compose file (GPU version) - docker-compose.yml

### Updated 2025/06/01
```
version: '3.8'
services:
  ollama:
    image: ollama/ollama:latest
    ports:
      - 11434:11434
    volumes:
      - .:/code
      - ./ollama/ollama:/root/.ollama
    container_name: ollama
    pull_policy: always
    tty: true
    restart: always
    networks:
      - ollama-docker
    deploy:
      resources:
        reservations:
          devices:
            - capabilities: [gpu]

  open-webui:
    image: ghcr.io/open-webui/open-webui:main
    container_name: open-webui
    depends_on:
      - ollama
    ports:
      - 8088:8080
    environment:
      - 'OLLAMA_API=http://ollama:11434/api'
    extra_hosts:
      - host.docker.internal:host-gateway
    restart: unless-stopped
    networks:
      - ollama-docker

networks:
  ollama-docker:
    external: false
```

## Start the Containers
```
docker compose up -d
```
## Visit localhost:8080

![image](https://blog.markkulab.net/content/markku/posts/build-your-ollama-ai-with-hardware-acceleration/images/1.png)

## Settings > Download and Install a Model
![image](https://blog.markkulab.net/content/markku/posts/build-your-ollama-ai-with-hardware-acceleration/images/2.png)

## Install the Linux Subsystem (Windows PowerShell)
```
wsl --version
wsl --update
wsl --install
wsl --list
```
![image](https://blog.markkulab.net/content/markku/posts/build-your-ollama-ai-with-hardware-acceleration/images/3.png)

## Open Ubuntu
![image](https://blog.markkulab.net/content/markku/posts/build-your-ollama-ai-with-hardware-acceleration/images/4.png)

## Run the Following Inside Ubuntu
Inside Ubuntu, run the following to set up the NVIDIA Container Toolkit:

```
sudo apt-get update
sudo apt-get install -y nvidia-cuda-toolkit nvidia-container-toolkit
```

## Configure Docker for GPU Support

     * Make sure Docker has **WSL2 integration enabled**
     * Open Docker Desktop → Settings → Resources → WSL integration → enable Ubuntu
     * Recent Docker versions Enable GPU support** (on by default) => no extra config needed
	 * Docker engine settings
![image](https://blog.markkulab.net/content/markku/posts/build-your-ollama-ai-with-hardware-acceleration/images/6.png)

```
"runtimes": {
    "nvidia": {
      "path": "nvidia-container-runtime",
      "runtimeArgs": []
    }
  }
```	 
## Test GPU Integration
Verify GPU integration works:
![image](https://blog.markkulab.net/content/markku/posts/build-your-ollama-ai-with-hardware-acceleration/images/5.png)
```
docker run --gpus all nvidia/cuda:11.5.2-base-ubuntu20.04 nvidia-smi
```

## Confirm Ollama Is Using the GPU — run the following on the host
```
docker exec -it ollama /bin/bash
ollama ps
```
![image](https://blog.markkulab.net/content/markku/posts/build-your-ollama-ai-with-hardware-acceleration/images/7.png)

## Benchmark Results
### Hardware: CPU 13900K + Nvidia TUF RTX 3080 + 64 GB RAM + Win 11
### We fed Ollama the same SEO-generation prompt we used with Azure OpenAI.

```
You are an SEO expert. Based on the page description provided below, generate an SEO-optimized title, meta description, and keywords. Ensure that the title is engaging and concise, the meta description summarizes the product effectively while enticing users to learn more, and the keywords are relevant to the product's features and market segment. Additionally, translate all content into the language specified by the given language code.
Company:Your company Desc ...
Ecommerce Page Description: Case: NZXT H5 Flow Gaming Gehäuse - Schwarz Processor: AMD Ryzen 5 5600X Processor (6x 3.7GHz/32MB L3 Cache) Memory: 16GB DDR4/3200MHz Memory(G.Skill ,Corsair,Kingston) Storage: Video Card: NVIDIA GeForce RTX 3050 - 8GB GDDR6X (VR-Ready) Motherboard: ASRock B450 PRO 4 ATX USB 3.1, SATA3, 1x M.2
Translate Target Language Code: en
FormatInstructions: Only the title, description and keywords of the json structure are returned. example :{"title":"","description":"","keywords":""} Please delete any other unnecessary information. Such as python code, Python Flask API, etc. Give me json result. Do not send back any other information such as python code, Python Flask API, etc.
```
### Performance Comparison
* CPU only: 1 min 10 sec ~ 1 min 30 sec
* NVIDIA RTX MSI 2060 OG GPU: around 30 seconds
* NVIDIA RTX TUF 3080 GPU: under 3 seconds

## References
* [Idiot's Guide to Hosting an LLM - Ollama + Open WebUI Docker Compose Setup](https://blog.darkthread.net/blog/ollam-open-webui/)
* [[CUDA] How to install Ollama 3 + Open WebUI on Windows (docker + WSL 2 + ubuntu + nvidia-container)](https://blog.csdn.net/smileyan9/article/details/140391667)
* [How to Use Ollama Elegantly | JD Cloud Tech Team](https://blog.csdn.net/jdcdev_/article/details/138800464)
* [A Collection of Common Ollama Models](https://www.53ai.com/news/qianyanjishu/1201.html)
* [No Privacy Worries Offline! Free Open-Source AI Assistant Ollama — From Install to Fine-Tuning in One Video](https://www.youtube.com/watch?v=JpQC0W91E6k&t=397s)

---

## About this article and its author

Originally published on [Mark Ku's Tech Notes](https://blog.markkulab.net/en/post/build-your-ollama-ai-with-hardware-acceleration)

License: [CC BY 4.0](https://creativecommons.org/licenses/by/4.0/) — when reusing or quoting, credit the author and link back to the original

### About the author

**[Mark Ku](https://blog.markkulab.net/en/author/mark-ku)** — Software Solution Provider

- 10+ years senior software engineer, now an AI Builder
- Focused on large-platform architecture — North-American e-commerce, AI SaaS subscription billing
- Combining AI Agents and automation to build evolvable product foundations

### Free tools built by the author

All of these are free to use:

- [Free PDF Sign Tool](https://blog.markkulab.net/en/tools/pdf-sign): Online PDF sign tool — draw, type, or upload a signature, then drag, resize, and download. Everything runs in your browser; nothing is uploaded.
- [VS Code Refactory](https://blog.markkulab.net/en/tools/refactory): Refactory is a VS Code refactoring extension: 34 actions plus a 37-rule code-smell inspection layer with a Code Health dashboard, across 18 languages, backed by 534 tests. It learns your repo's conventions: where interfaces live, where DI is registered, whether 'use client' belongs. It ranks files by git churn × complexity so you know what to fix first, and hands any smell to the Claude Code already on your machine. Free to use, and your source never leaves your computer.
- [DB-Kit Database Manager](https://blog.markkulab.net/en/tools/db-kit): DB-Kit is a lightweight, cross-platform database manager built with Tauri + Rust + React. Manage MySQL, MariaDB, PostgreSQL, SQL Server, Oracle, SQLite, MongoDB, Redis, Kafka, Elasticsearch and RabbitMQ from one consistent interface: passwords encrypted in the OS keychain, SSH tunnels, full CRUD, a visual query builder, stacked multi-statement result sets, cross-connection data transfer and compare/sync, Excel / CSV import & export, visualized execution plans, ER diagrams, scheduled backups, SQL stress testing with p50–p99 latency percentiles, a 15-rule SQL review engine, Kafka message browsing with monitoring & alerts, a bilingual UI (Traditional Chinese / English), a built-in AI assistant (natural-language SQL, AI review and tuning advice) and the dbk CLI. Free and open source (MIT), with installers for Windows, macOS and Linux.
- [VS Code Super Mermaid](https://blog.markkulab.net/en/tools/super-mermaid): Super Mermaid is a VS Code extension for beautiful Mermaid diagrams out of the box: auto-colored live preview, mouse pan & zoom, high-res PNG / SVG export, 21 templates and multiple themes. Free and open source (MIT).
- [React Super Mermaid](https://blog.markkulab.net/en/tools/react-super-mermaid): react-super-mermaid is an open-source React component library: render beautiful Mermaid diagrams with a single <MermaidViewer>, with built-in colorful / sketch themes, pan & zoom, in-diagram search, and high-res SVG / PNG export. Lightweight, SSR-safe, fully typed. Free and open source (MIT).
- [Jira / Confluence Super Mermaid](https://blog.markkulab.net/en/tools/jira-super-mermaid): An Atlassian Forge app: write Mermaid syntax directly inside a Jira issue or a Confluence page and get flowcharts, sequence diagrams, state machines and Gantt charts. 11 diagram types, SVG / PNG export, light and dark themes, full CJK support. Runs on Atlassian: your diagrams live in your own site and the app calls no third-party service. Free, coming soon to the Atlassian Marketplace.
- [Mermaid Live Preview](https://blog.markkulab.net/en/tools/mermaid-preview): Write Mermaid in your browser, see it render instantly, and share the whole diagram as a single link. No sign-up, nothing uploaded to a server, and mermaid.live share links work as-is.
- [React Intl Phone Number](https://blog.markkulab.net/en/tools/react-intl-phone-number): react-intl-phone-number is an open-source React component: framework-agnostic and antd-free, with E.164 in/out, a searchable flag / country-code dropdown, configurable validation levels (strict / mobile-strict / loose), themeable CSS, and i18n — phone logic powered by google-libphonenumber. Lightweight and fully typed. Free and open source (MIT).
- [Uptime Kuma Cluster](https://blog.markkulab.net/en/tools/uptime-kuma-cluster): Turn single-node Uptime Kuma into a highly available cluster: OpenResty + Lua smart load balancing, shared MariaDB state, health checks and automatic failover, plus cluster-management REST APIs. One Docker Compose command to start. Free and open source (MIT).
- [Special Education](https://blog.markkulab.net/en/education): Learning materials crafted for special education students

### Daily podcasts

- [Mark's Tech Insights — Daily AI News](https://blog.markkulab.net/en/category/tech-news): Daily curated AI and tech trends. Catch the latest developments via audio summaries — covering AI applications, software architecture, DevOps, and engineering practice. — RSS: https://blog.markkulab.net/feed.xml
- [AI股市蝦聊](https://blog.markkulab.net/en/category/ai-stock-chat): Every trading day, an AI-analyzed take on the Taiwan stock market, delivered as a two-host conversation covering the session and the next-day outlook. — RSS: https://blog.markkulab.net/ai-stock-chat/feed.xml
- [開源好物週報](https://blog.markkulab.net/en/category/open-source-weekly): A weekly two-host pick of free open-source tools surfaced from real Hacker News, GitHub, and Reddit buzz — what pain they solve and the fastest way to get started. — RSS: https://blog.markkulab.net/open-source-weekly/feed.xml

### Newsletter

[Subscribe to the newsletter](https://blog.markkulab.net/en/subscribe) — Be the first to know about new posts. No spam, unsubscribe anytime.
