Private by default
We run no chat server. Your conversations are processed and stored in your browser.
Open source · Free · English & French
WebSLM runs open small language models right inside your browser. There is nothing to install and no account to create, and nothing you type is sent to a server.
Privacy check · this message contains a Social Insurance Number. It stays on this device.
Why WebSLM
The reasoning behind the project, in six short parts.
Using an AI assistant today usually means sending your text to a company's servers. For a grocery list that's fine. For a client file, a medical note, a tenant's letter or an employee review, it's a step many people shouldn't take, and many workplaces don't allow it.
Small language models, from a few hundred million to a few billion parameters, are now genuinely useful for everyday language work: rewriting, summarizing, translating and drafting. They are small enough to download once and run on an ordinary laptop.
WebAssembly and WebGPU let a web page run demanding code at close to native speed. llama.cpp, the open-source engine behind much of local AI, now runs in the browser. A website can run a model without the website ever seeing your data.
Open a page, pick a model, and chat. The model runs on your machine. There is nothing to install and no account, and after the first visit it works offline. Because the code is open source, you don't have to take our word for any of this: you can read it.
The whole interface works in English and French, with Canadian French translation built in. The privacy check recognizes Canadian personal information, such as Social Insurance Numbers (validated by checksum), health card numbers and postal codes, and offers to mask it before you send.
Small models are not the largest cloud models. They make mistakes and know less about the world, so check any facts they give you. They are the right tool for a large share of everyday writing. When you need more, WebSLM connects to a bigger model on your own llama-server, which stays under your control.
We run no chat server. Your conversations are processed and stored in your browser.
A link is all it takes: no installer, no account and no API key.
Apache-2.0 code, open models and clear licences shown next to every model.
English and French interface, tasks and privacy check, not an afterthought.
How it works
Pick one of five curated open models, from about 230 MB to 1 GB. Each shows its size and licence.
Your browser fetches the model file from Hugging Face and keeps it in its cache. Next time it loads from disk.
llama.cpp runs the model on your device, using WebGPU when available and your CPU otherwise.
Run llama.cpp's server on your own computer and switch engines in one click. Messages then go only to the address you choose.
llama-server -hf LiquidAI/LFM2.5-1.2B-Instruct-GGUF --port 8080
Rewrite dense text at a grade 8 reading level without losing the meaning.
English to Canadian French and back, with only the translation returned.
Up to five bullet points, plus a line listing any action items.
Turn rough notes into a short, polite email with a subject line.
Privacy check
Even on your own device, it helps to know what's in a message, especially before you connect to a server or share a chat. WebSLM scans each message as you send it and lets you mask what it finds. Exports to Markdown can mask it too.
Models
Each model is downloaded directly from its publisher's page on Hugging Face. WebSLM doesn't bundle or modify them.
| Model | Download | Licence | Best for |
|---|
Before you start
WebSLM is free, and it runs open models made by other organizations. That comes with a few things you should know before you rely on it.
WebSLM is free open-source software under the Apache-2.0 licence, provided "as is", without warranties of any kind. There is no paid support or uptime guarantee.
Small models can produce confident text that is incorrect, outdated, incomplete or biased. Review everything before you rely on it or share it.
Don't use WebSLM in place of medical, legal, financial or other professional advice, or to make decisions that significantly affect people.
The models are made by Liquid AI, IBM, Alibaba Qwen and Hugging Face, not by WebSLM. You are responsible for following each model's licence and usage policy. For example, Liquid AI's LFM licence is free only for organizations under US$10M in annual revenue.
It looks for common patterns. It can miss personal information or flag text that isn't personal. Read your message before you send or share it.
A model takes about 230 MB to 1 GB of storage and more memory while it runs. Generating text keeps your processor or graphics card busy and drains battery faster. Use Wi-Fi for the first download if your data plan is limited. You can remove models at any time in the app.
Conversations stay in your browser's storage until you start a new chat or clear site data. On a shared computer, start a new chat when you're done.
Loading this site contacts GitHub Pages, and downloading a model contacts Hugging Face. Like any website, they can see your IP address, and their own privacy policies apply. Anonymous visit and download counts go to GoatCounter. Your prompts are not sent to any of them.
Your prompts and the model's answers are computed in your browser and never sent to us, because we don't run a chat server at all. Only two things use the network: loading this website (hosted on GitHub Pages) and the one-time model download from Hugging Face. If you connect your own llama-server, messages go only to the address you enter.
We count anonymous visits to these pages and model downloads (which page, which model, and a coarse device class: WebGPU or CPU, and low, mid or high capability) with GoatCounter, which sets no cookies and doesn't store IP addresses. We never count or collect anything you type. Browsers that send Do Not Track or Global Privacy Control aren't counted at all.
Nothing. WebSLM is free and open source under the Apache-2.0 licence.
A recent desktop browser. Browsers with WebGPU are fastest; without it, WebSLM uses your CPU through WebAssembly. The first model download is between about 230 MB and 1 GB.
No, and it isn't trying to be. Small models are best at rewriting, summarizing, translating and drafting. They can be wrong, so check facts. For heavier work, connect a larger model on your own llama-server.
The WebSLM code is Apache-2.0. Each model has its own licence, shown next to it in the app. For example, Liquid AI's LFM models are free for organizations under US$10M in annual revenue, while Granite, Qwen3 and SmolLM2 are Apache-2.0. Check the licence and your organization's policies.
The app checks your device first: whether WebGPU is available, roughly how much memory your browser reports, your CPU threads and your free storage. It suggests a model and marks each one as a good fit, possibly slow, or too big for your free storage. These checks happen on your device. As a rule of thumb, start with Liquid LFM2.5 1.2B on a recent laptop and with LFM2.5 350M or SmolLM2 360M on older machines and phones.
In the app, under Model, each downloaded model shows its size and a Remove link. The Your device panel also has Remove all downloaded models. Removing a model frees the space right away and keeps your chats. You can also clear this site's data in your browser settings, which removes models and chats.
Yes. After your first visit and model download, the app and the model are cached by your browser, so you can keep chatting without a connection.