Getting Started with WebLLM
Learn how to run large language models directly in the browser using WebGPU. No server required!
prerequisites
2- Basic JavaScript knowledge
- Familiarity with async/await
WebLLM enables running large language models directly in your browser using WebGPU acceleration. This means you can build AI-powered applications without sending data to external servers—everything runs locally on the user's device.
Prerequisites
Before we begin, make sure you have:
- A browser that supports WebGPU (Chrome 113+ or Edge 113+)
- Basic knowledge of JavaScript and async/await
- Node.js installed for the development setup
What is WebLLM?
WebLLM is a library developed by the MLC AI team that brings large language models to the browser. It uses:
- WebGPU: A new web standard for GPU-accelerated computation
- MLC-LLM: Machine Learning Compilation for efficient model execution
- Quantized Models: Compressed model weights that fit in browser memory
Step 1: Setting Up Your Project
First, create a new project and install WebLLM:
# Create a new Vite project
npm create vite@latest webllm-demo -- --template vanilla-ts
cd webllm-demo
# Install WebLLM
npm install @mlc-ai/web-llmStep 2: Basic WebLLM Setup
Create a simple chat interface. In your main.ts:
import * as webllm from "@mlc-ai/web-llm";
// Available models - smaller models load faster
const MODEL_ID = "SmolLM2-360M-Instruct-q4f16_1-MLC";
class WebLLMChat {
private engine: webllm.MLCEngine | null = null;
async initialize(onProgress: (progress: string) => void) {
// Create the engine
this.engine = new webllm.MLCEngine();
// Set up progress callback
this.engine.setInitProgressCallback((report) => {
onProgress(`Loading: ${report.text} (${Math.round(report.progress * 100)}%)`);
});
// Load the model
await this.engine.reload(MODEL_ID);
onProgress("Ready!");
}
async chat(message: string): Promise<string> {
if (!this.engine) throw new Error("Engine not initialized");
const response = await this.engine.chat.completions.create({
messages: [{ role: "user", content: message }],
temperature: 0.7,
max_tokens: 256,
});
return response.choices[0].message.content || "";
}
}
// Usage
const chat = new WebLLMChat();
document.querySelector<HTMLDivElement>("#app")!.innerHTML = `
<div>
<h1>WebLLM Chat</h1>
<div id="status">Initializing...</div>
<input type="text" id="input" placeholder="Type a message..." disabled />
<button id="send" disabled>Send</button>
<div id="response"></div>
</div>
`;
const statusEl = document.getElementById("status")!;
const inputEl = document.getElementById("input") as HTMLInputElement;
const sendBtn = document.getElementById("send") as HTMLButtonElement;
const responseEl = document.getElementById("response")!;
// Initialize
chat.initialize((status) => {
statusEl.textContent = status;
if (status === "Ready!") {
inputEl.disabled = false;
sendBtn.disabled = false;
}
});
// Handle send
sendBtn.addEventListener("click", async () => {
const message = inputEl.value;
if (!message) return;
sendBtn.disabled = true;
responseEl.textContent = "Thinking...";
const response = await chat.chat(message);
responseEl.textContent = response;
sendBtn.disabled = false;
inputEl.value = "";
});Step 3: Adding Streaming Responses
For a better user experience, stream the response token by token:
async chatStream(
message: string,
onToken: (token: string) => void
): Promise<void> {
if (!this.engine) throw new Error("Engine not initialized");
const response = await this.engine.chat.completions.create({
messages: [{ role: "user", content: message }],
temperature: 0.7,
max_tokens: 256,
stream: true, // Enable streaming
});
// Iterate over the stream
for await (const chunk of response) {
const token = chunk.choices[0]?.delta?.content || "";
if (token) {
onToken(token);
}
}
}
// Usage with streaming
let fullResponse = "";
await chat.chatStream(message, (token) => {
fullResponse += token;
responseEl.textContent = fullResponse;
});Step 4: Handling Model Loading
Model loading can take time (downloading ~200MB-2GB depending on the model). Here's how to provide good feedback:
interface LoadingState {
stage: "downloading" | "caching" | "initializing" | "ready";
progress: number;
text: string;
}
function parseProgress(report: webllm.InitProgressReport): LoadingState {
const text = report.text.toLowerCase();
if (text.includes("fetch")) {
return {
stage: "downloading",
progress: report.progress,
text: "Downloading model weights..."
};
}
if (text.includes("cache")) {
return {
stage: "caching",
progress: report.progress,
text: "Caching model for faster loads..."
};
}
if (text.includes("loading") || text.includes("init")) {
return {
stage: "initializing",
progress: report.progress,
text: "Initializing model..."
};
}
return {
stage: "ready",
progress: 1,
text: "Ready!"
};
}Step 5: Conversation Memory
For multi-turn conversations, maintain message history:
interface Message {
role: "user" | "assistant" | "system";
content: string;
}
class ConversationalChat {
private engine: webllm.MLCEngine | null = null;
private messages: Message[] = [];
// Add system prompt
setSystemPrompt(prompt: string) {
this.messages = [{ role: "system", content: prompt }];
}
async chat(userMessage: string): Promise<string> {
if (!this.engine) throw new Error("Engine not initialized");
// Add user message to history
this.messages.push({ role: "user", content: userMessage });
// Get response with full history
const response = await this.engine.chat.completions.create({
messages: this.messages,
temperature: 0.7,
max_tokens: 256,
});
const assistantMessage = response.choices[0].message.content || "";
// Add assistant response to history
this.messages.push({ role: "assistant", content: assistantMessage });
return assistantMessage;
}
clearHistory() {
// Keep system prompt if present
this.messages = this.messages.filter(m => m.role === "system");
}
}Available Models
WebLLM supports various models. Smaller models are faster but less capable:
| Model | Size | Use Case |
|---|---|---|
| SmolLM2-360M-Instruct | ~200MB | Quick responses, simple tasks |
| Llama-3.2-1B-Instruct | ~600MB | Better quality, still fast |
| Llama-3.2-3B-Instruct | ~1.5GB | Good balance of speed/quality |
| Phi-3.5-mini-instruct | ~2GB | Strong reasoning |
Common Issues
WebGPU Not Available
async function checkWebGPU(): Promise<boolean> {
if (!navigator.gpu) {
console.error("WebGPU not supported in this browser");
return false;
}
const adapter = await navigator.gpu.requestAdapter();
if (!adapter) {
console.error("No GPU adapter found");
return false;
}
return true;
}Memory Issues
If the model fails to load, try a smaller model or check available GPU memory.
Next Steps
Now that you have basic WebLLM working:
- Add a proper UI: Build a chat interface with message history
- Implement error handling: Handle network issues and GPU errors gracefully
- Explore other models: Try different models for your use case
- Add features: Implement copy-to-clipboard, message editing, etc.
In the next tutorial, we'll build a complete React chat interface with WebLLM.
This tutorial is part of the WebLLM Fundamentals series.