Open source · Free · English & French

AI that stays on your device.

WebSLM runs open small language models right inside your browser. There is nothing to install and no account to create, and nothing you type is sent to a server.

  • On-deviceprompts never leave your computer
  • Offlineafter the first visit
  • Apache-2.0read every line of code
webslm.org/app On this device
Summarize this note: J. Tremblay (SIN 046 454 286) called about her renewal and asked for a callback Friday.

Privacy check · this message contains a Social Insurance Number. It stays on this device.

Mask and send Send anyway
• Client J. Tremblay called about her renewal.
• Action: call her back on Friday.
Liquid LFM2.5 1.2B · WebGPU
Illustration. Personal information is flagged before you send.

Why WebSLM

Useful AI shouldn't require giving away your words.

The reasoning behind the project, in six short parts.

  1. Most AI asks you to hand over what you write.

    Using an AI assistant today usually means sending your text to a company's servers. For a grocery list that's fine. For a client file, a medical note, a tenant's letter or an employee review, it's a step many people shouldn't take, and many workplaces don't allow it.

  2. Small models became good enough.

    Small language models, from a few hundred million to a few billion parameters, are now genuinely useful for everyday language work: rewriting, summarizing, translating and drafting. They are small enough to download once and run on an ordinary laptop.

  3. The browser became a real computer.

    WebAssembly and WebGPU let a web page run demanding code at close to native speed. llama.cpp, the open-source engine behind much of local AI, now runs in the browser. A website can run a model without the website ever seeing your data.

  4. So we put the two together.

    Open a page, pick a model, and chat. The model runs on your machine. There is nothing to install and no account, and after the first visit it works offline. Because the code is open source, you don't have to take our word for any of this: you can read it.

  5. Made with Canada in mind.

    The whole interface works in English and French, with Canadian French translation built in. The privacy check recognizes Canadian personal information, such as Social Insurance Numbers (validated by checksum), health card numbers and postal codes, and offers to mask it before you send.

  6. Honest about the limits.

    Small models are not the largest cloud models. They make mistakes and know less about the world, so check any facts they give you. They are the right tool for a large share of everyday writing. When you need more, WebSLM connects to a bigger model on your own llama-server, which stays under your control.

What we hold ourselves to

Private by default

We run no chat server. Your conversations are processed and stored in your browser.

Zero setup

A link is all it takes: no installer, no account and no API key.

Open and verifiable

Apache-2.0 code, open models and clear licences shown next to every model.

Bilingual from day one

English and French interface, tasks and privacy check, not an afterthought.

How it works

Three steps, all on your machine

  1. 1

    Choose a model

    Pick one of five curated open models, from about 230 MB to 1 GB. Each shows its size and licence.

  2. 2

    It downloads once

    Your browser fetches the model file from Hugging Face and keeps it in its cache. Next time it loads from disk.

  3. 3

    Chat privately

    llama.cpp runs the model on your device, using WebGPU when available and your CPU otherwise.

Need more speed or a bigger model?

Run llama.cpp's server on your own computer and switch engines in one click. Messages then go only to the address you choose.

llama-server -hf LiquidAI/LFM2.5-1.2B-Instruct-GGUF --port 8080

Built-in tasks for everyday writing

Plain language

Rewrite dense text at a grade 8 reading level without losing the meaning.

Translate EN ↔ FR

English to Canadian French and back, with only the translation returned.

Summarize

Up to five bullet points, plus a line listing any action items.

Draft an email

Turn rough notes into a short, polite email with a subject line.

Privacy check

A second look before anything sensitive goes out

Even on your own device, it helps to know what's in a message, especially before you connect to a server or share a chat. WebSLM scans each message as you send it and lets you mask what it finds. Exports to Markdown can mask it too.

  • Social Insurance Number (checksum)
  • Health card numbers: AB, ON, QC
  • Passport numbers
  • Business numbers
  • Postal codes
  • Email addresses
  • Phone numbers
  • Payment card numbers

Models

Curated small models, clearly licensed

Each model is downloaded directly from its publisher's page on Hugging Face. WebSLM doesn't bundle or modify them.

ModelDownloadLicenceBest for

Before you start

Good to know: free software, open models

WebSLM is free, and it runs open models made by other organizations. That comes with a few things you should know before you rely on it.

Free, and provided as is

WebSLM is free open-source software under the Apache-2.0 licence, provided "as is", without warranties of any kind. There is no paid support or uptime guarantee.

Answers can be wrong

Small models can produce confident text that is incorrect, outdated, incomplete or biased. Review everything before you rely on it or share it.

Not professional advice

Don't use WebSLM in place of medical, legal, financial or other professional advice, or to make decisions that significantly affect people.

Each model has its own licence

The models are made by Liquid AI, IBM, Alibaba Qwen and Hugging Face, not by WebSLM. You are responsible for following each model's licence and usage policy. For example, Liquid AI's LFM licence is free only for organizations under US$10M in annual revenue.

The privacy check is a helper

It looks for common patterns. It can miss personal information or flag text that isn't personal. Read your message before you send or share it.

Your device does the work

A model takes about 230 MB to 1 GB of storage and more memory while it runs. Generating text keeps your processor or graphics card busy and drains battery faster. Use Wi-Fi for the first download if your data plan is limited. You can remove models at any time in the app.

Chats are saved in this browser

Conversations stay in your browser's storage until you start a new chat or clear site data. On a shared computer, start a new chat when you're done.

Other services you connect to

Loading this site contacts GitHub Pages, and downloading a model contacts Hugging Face. Like any website, they can see your IP address, and their own privacy policies apply. Anonymous visit and download counts go to GoatCounter. Your prompts are not sent to any of them.

Questions

Is it really private?

Your prompts and the model's answers are computed in your browser and never sent to us, because we don't run a chat server at all. Only two things use the network: loading this website (hosted on GitHub Pages) and the one-time model download from Hugging Face. If you connect your own llama-server, messages go only to the address you enter.

Do you track me?

We count anonymous visits to these pages and model downloads (which page, which model, and a coarse device class: WebGPU or CPU, and low, mid or high capability) with GoatCounter, which sets no cookies and doesn't store IP addresses. We never count or collect anything you type. Browsers that send Do Not Track or Global Privacy Control aren't counted at all.

What does it cost?

Nothing. WebSLM is free and open source under the Apache-2.0 licence.

What do I need?

A recent desktop browser. Browsers with WebGPU are fastest; without it, WebSLM uses your CPU through WebAssembly. The first model download is between about 230 MB and 1 GB.

Is it as capable as the big cloud assistants?

No, and it isn't trying to be. Small models are best at rewriting, summarizing, translating and drafting. They can be wrong, so check facts. For heavier work, connect a larger model on your own llama-server.

Can I use it at work?

The WebSLM code is Apache-2.0. Each model has its own licence, shown next to it in the app. For example, Liquid AI's LFM models are free for organizations under US$10M in annual revenue, while Granite, Qwen3 and SmolLM2 are Apache-2.0. Check the licence and your organization's policies.

Which model should I download?

The app checks your device first: whether WebGPU is available, roughly how much memory your browser reports, your CPU threads and your free storage. It suggests a model and marks each one as a good fit, possibly slow, or too big for your free storage. These checks happen on your device. As a rule of thumb, start with Liquid LFM2.5 1.2B on a recent laptop and with LFM2.5 350M or SmolLM2 360M on older machines and phones.

How do I remove a downloaded model?

In the app, under Model, each downloaded model shows its size and a Remove link. The Your device panel also has Remove all downloaded models. Removing a model frees the space right away and keeps your chats. You can also clear this site's data in your browser settings, which removes models and chats.

Does it work offline?

Yes. After your first visit and model download, the app and the model are cached by your browser, so you can keep chatting without a connection.

Try it now. It takes one download.

No sign-up, and your words stay on your device.

Open WebSLM