Blogsbeginner

Getting Started with WebLLM

Learn how to run large language models directly in the browser using WebGPU. No server required!

15 minWebLLM, WebGPU, JavaScriptSeries: WebLLM Fundamentals

prerequisites

2
  • Basic JavaScript knowledge
  • Familiarity with async/await

WebLLM enables running large language models directly in your browser using WebGPU acceleration. This means you can build AI-powered applications without sending data to external servers—everything runs locally on the user's device.

Prerequisites

Before we begin, make sure you have:

  • A browser that supports WebGPU (Chrome 113+ or Edge 113+)
  • Basic knowledge of JavaScript and async/await
  • Node.js installed for the development setup

What is WebLLM?

WebLLM is a library developed by the MLC AI team that brings large language models to the browser. It uses:

  • WebGPU: A new web standard for GPU-accelerated computation
  • MLC-LLM: Machine Learning Compilation for efficient model execution
  • Quantized Models: Compressed model weights that fit in browser memory

Step 1: Setting Up Your Project

First, create a new project and install WebLLM:

Bash
# Create a new Vite project
npm create vite@latest webllm-demo -- --template vanilla-ts
cd webllm-demo

# Install WebLLM
npm install @mlc-ai/web-llm

Step 2: Basic WebLLM Setup

Create a simple chat interface. In your main.ts:

TypeScript
import * as webllm from "@mlc-ai/web-llm";

// Available models - smaller models load faster
const MODEL_ID = "SmolLM2-360M-Instruct-q4f16_1-MLC";

class WebLLMChat {
  private engine: webllm.MLCEngine | null = null;

  async initialize(onProgress: (progress: string) => void) {
    // Create the engine
    this.engine = new webllm.MLCEngine();

    // Set up progress callback
    this.engine.setInitProgressCallback((report) => {
      onProgress(`Loading: ${report.text} (${Math.round(report.progress * 100)}%)`);
    });

    // Load the model
    await this.engine.reload(MODEL_ID);
    onProgress("Ready!");
  }

  async chat(message: string): Promise<string> {
    if (!this.engine) throw new Error("Engine not initialized");

    const response = await this.engine.chat.completions.create({
      messages: [{ role: "user", content: message }],
      temperature: 0.7,
      max_tokens: 256,
    });

    return response.choices[0].message.content || "";
  }
}

// Usage
const chat = new WebLLMChat();

document.querySelector<HTMLDivElement>("#app")!.innerHTML = `
  <div>
    <h1>WebLLM Chat</h1>
    <div id="status">Initializing...</div>
    <input type="text" id="input" placeholder="Type a message..." disabled />
    <button id="send" disabled>Send</button>
    <div id="response"></div>
  </div>
`;

const statusEl = document.getElementById("status")!;
const inputEl = document.getElementById("input") as HTMLInputElement;
const sendBtn = document.getElementById("send") as HTMLButtonElement;
const responseEl = document.getElementById("response")!;

// Initialize
chat.initialize((status) => {
  statusEl.textContent = status;
  if (status === "Ready!") {
    inputEl.disabled = false;
    sendBtn.disabled = false;
  }
});

// Handle send
sendBtn.addEventListener("click", async () => {
  const message = inputEl.value;
  if (!message) return;

  sendBtn.disabled = true;
  responseEl.textContent = "Thinking...";

  const response = await chat.chat(message);
  responseEl.textContent = response;

  sendBtn.disabled = false;
  inputEl.value = "";
});

Step 3: Adding Streaming Responses

For a better user experience, stream the response token by token:

TypeScript
async chatStream(
  message: string,
  onToken: (token: string) => void
): Promise<void> {
  if (!this.engine) throw new Error("Engine not initialized");

  const response = await this.engine.chat.completions.create({
    messages: [{ role: "user", content: message }],
    temperature: 0.7,
    max_tokens: 256,
    stream: true,  // Enable streaming
  });

  // Iterate over the stream
  for await (const chunk of response) {
    const token = chunk.choices[0]?.delta?.content || "";
    if (token) {
      onToken(token);
    }
  }
}

// Usage with streaming
let fullResponse = "";
await chat.chatStream(message, (token) => {
  fullResponse += token;
  responseEl.textContent = fullResponse;
});

Step 4: Handling Model Loading

Model loading can take time (downloading ~200MB-2GB depending on the model). Here's how to provide good feedback:

TypeScript
interface LoadingState {
  stage: "downloading" | "caching" | "initializing" | "ready";
  progress: number;
  text: string;
}

function parseProgress(report: webllm.InitProgressReport): LoadingState {
  const text = report.text.toLowerCase();

  if (text.includes("fetch")) {
    return {
      stage: "downloading",
      progress: report.progress,
      text: "Downloading model weights..."
    };
  }

  if (text.includes("cache")) {
    return {
      stage: "caching",
      progress: report.progress,
      text: "Caching model for faster loads..."
    };
  }

  if (text.includes("loading") || text.includes("init")) {
    return {
      stage: "initializing",
      progress: report.progress,
      text: "Initializing model..."
    };
  }

  return {
    stage: "ready",
    progress: 1,
    text: "Ready!"
  };
}

Step 5: Conversation Memory

For multi-turn conversations, maintain message history:

TypeScript
interface Message {
  role: "user" | "assistant" | "system";
  content: string;
}

class ConversationalChat {
  private engine: webllm.MLCEngine | null = null;
  private messages: Message[] = [];

  // Add system prompt
  setSystemPrompt(prompt: string) {
    this.messages = [{ role: "system", content: prompt }];
  }

  async chat(userMessage: string): Promise<string> {
    if (!this.engine) throw new Error("Engine not initialized");

    // Add user message to history
    this.messages.push({ role: "user", content: userMessage });

    // Get response with full history
    const response = await this.engine.chat.completions.create({
      messages: this.messages,
      temperature: 0.7,
      max_tokens: 256,
    });

    const assistantMessage = response.choices[0].message.content || "";

    // Add assistant response to history
    this.messages.push({ role: "assistant", content: assistantMessage });

    return assistantMessage;
  }

  clearHistory() {
    // Keep system prompt if present
    this.messages = this.messages.filter(m => m.role === "system");
  }
}

Available Models

WebLLM supports various models. Smaller models are faster but less capable:

ModelSizeUse Case
SmolLM2-360M-Instruct~200MBQuick responses, simple tasks
Llama-3.2-1B-Instruct~600MBBetter quality, still fast
Llama-3.2-3B-Instruct~1.5GBGood balance of speed/quality
Phi-3.5-mini-instruct~2GBStrong reasoning

Common Issues

WebGPU Not Available

TypeScript
async function checkWebGPU(): Promise<boolean> {
  if (!navigator.gpu) {
    console.error("WebGPU not supported in this browser");
    return false;
  }

  const adapter = await navigator.gpu.requestAdapter();
  if (!adapter) {
    console.error("No GPU adapter found");
    return false;
  }

  return true;
}

Memory Issues

If the model fails to load, try a smaller model or check available GPU memory.

Next Steps

Now that you have basic WebLLM working:

  1. Add a proper UI: Build a chat interface with message history
  2. Implement error handling: Handle network issues and GPU errors gracefully
  3. Explore other models: Try different models for your use case
  4. Add features: Implement copy-to-clipboard, message editing, etc.

In the next tutorial, we'll build a complete React chat interface with WebLLM.


This tutorial is part of the WebLLM Fundamentals series.

Zizhao Huloading加载中