v0.1.0No server, no API key, nothing leaves your browser

Semantic search is now open, fast, free and available for everyone

in your browser & Node.js for fast search < 100 ms for ~97% recall @ top 5 with just 39.8 MB model all-in and ~21,5 kB index for 100 docs who need fast, efficient semantic search

The Winzling embedding model is an extremely small text embedding model created with specific quantization and pruning methods. With a model size of 39.8 MB all-in it achieves ~97% recall in German, English and other languages at top 5 - usually within < 100ms even on the weakest devices and in-browser.

Model, inference and vector search algorithms brought to you by Aron Homberg (kyr0) (opens in a new tab) – Open Source (MIT (opens in a new tab))

Small enough to ship

4-bit TurboQuant stores each document in 256 bytes, so a whole index travels to the client like an image does. Vectors have 384 dimensions; a text may run to 8,192 tokens.

Runs on modest hardware

One codebase for Node.js and the browser, on CPU/WASM with no GPU: < 100 ms end to end and about 216 MB RAM. Made for old phones, weak business laptops and embedded boards.

Open Research, Customizable

make bench regenerates every figure on this page from a tiny benchmark (opens in a new tab). The method carries over to other languages, multimodal, late-interaction and fine-tuned embedders. Do you need a custom model? Ask Aron Homberg (LinkedIn) (opens in a new tab).

Zero cost - it's free!

MIT-licensed code, commercially usable, public model files on Hugging Face and static hosting: no API keys, no per-query fees, no servers to run. This page is plain files.

Perfect when privacy matters

Queries and your own documents never leave the device: no server, no third-party AI API, no data transfer to justify under GDPR (DSGVO). Only the model and runtime files are fetched; self-host them to keep even that in-house.

Pareto-optimal quality

Good enough for most retrieval and ranking: the right passage is in the top 5 for ~97% of German, English and Russian queries (92.2% across all nine target languages) and in the top 25 for 98.9%, at 39.8 MB instead of 305 MB for the larger Harrier. Measured on 2,000 passages.

Tokenize.

Embed.

Retrieve.

Live demo

🔍 Finding a needle in a haystack

100 Belebele passages, each in 20 languages: 2,000 passages in one index. Pick one of the 1,000 example queries with ⌘K, or type your own. Or create your own vector database index right in your browser, in the tab.

Runtime

Loading the index. The model loads with your first search.

index
Index · 2,000 passages
vector index and passage text, gzipped
Model · Winzling uint4
Runtime · ONNX Runtime Web
~216 MB RAM reserved for model, runtime and index (measured in Chromium)

Backend

–

ONNX Runtime Web, 1 thread

Ready in

–

download, compile, warm-up

Vector Database Index

–

documents

Last query

–

embed + TurboQuant scan

Search 2,000 passages in 20 languages

Enter a question above and press Search, or pick one of the 1,000 example queries. The first search downloads the 39.8 MB model once; every search after that runs on this device as you type.

    How it works

    💡 One query, three steps, all on your device

    Documents are embedded once, ahead of search: this demo's passages in Node.js, your own notes right in the browser. A search embeds only the query and compares it with the compressed vectors.

    Index · Node.js or browser
    Search · browser or Node.js
    01 · Text
    02 · Embed
    03 · TurboQuant
    04 · Result
    Documents 2,000 passages 20 languages
    Winzling Embed once Node.js or browser · 384-d
    TurboQuant 4-bit codes 256 B per passage
    Vector index 431 kB · file or memory
    User search query any language
    Winzling Embed the query ~21 ms in Chromium · 64k tokenizer
    • 39.8 MB, cached
    TurboQuant Scan 2,000 codes 0.4 ms
    Top 12, ranked by meaning
    1. text
    2. vectors
    3. codes
    4. 431 kB · once
    5. text
    6. 384-d
    7. top k

    Measured, not promised*

    Winzling in headless Chromium (WASM, 1 thread) on the same 2,000-passage haystack, from make bench. Winzling-Embed-a8m-64k (opens in a new tab) is a pruned model that is further pruned for a few select languages, and then custom quantized to extremely shrink it in size. Therefore, low performance in non-target languages is by design.

    21.3 msto embed a query (p50)
    0.4 msto scan the TurboQuant index (p50)
    431 kB gzipTurboQuant index for 2,000 passages; 0.51 MB raw, 6× smaller than float32
    100%R@5 for English queriesAll languages (opens in a new tab)

    * measured on an Apple MacBook Air M4 in a Chromium-based browser.

    R@5 by query language

    Share of a language's 50 questions whose gold passage ranks in the top 5 of 2,000. Chromium, TurboQuant. Target: the nine languages Winzling is built for. Pruned: below 70% R@5.

    English target
    German target
    Russian target
    French target
    Indonesian target
    Portuguese target
    Italian target
    Dutch target
    Spanish target
    Polish
    Turkish
    Vietnamese Pruned
    Korean Pruned
    Japanese Pruned
    Chinese Pruned
    Thai Pruned
    Urdu Pruned
    Hindi Pruned
    Bangla Pruned
    Arabic Pruned
    Portrait of Aron Homberg

    Need it tuned to your languages and data?

    Aron Homberg, freelance and independent AI researcher: custom embedding models, vector search and on-device AI, offers consulting services worldwide.

    Use it

    🧑‍💻 Add it to your own app

    Write the six lines yourself, or hand one prompt to your coding agent.

    NPM package

    6 lines in your frontend app. The first call downloads the pinned model from Hugging Face, checks every file's SHA-256 and caches it in the browser.

    search.ts
    import { buildTurboQuantIndex, createWinzlingEmbedder, searchTurboQuantIndex } from "defuss-vectorsearch/browser.js";const embedder = createWinzlingEmbedder();const documents = ["The cat sleeps on the sofa.", "Die Börse schloss heute im Plus.", "Кошка спит на диване."];const index = buildTurboQuantIndex(await embedder.embedDocuments(documents));const hits = searchTurboQuantIndex(index, await embedder.embedQuery("A sleeping cat on a couch."), 2);console.log(hits.map((hit) => documents[hit.index])); // both cat sentences

    Vibe coding

    Paste this prompt into your coding agent. It links the repository, so the agent works from the current README.

    Integrate defuss-vectorsearch into this app: https://github.com/kyr0/defuss-vectorsearch
    Read its README first and follow the browser quick start: import from "defuss-vectorsearch/browser.js", embed with createWinzlingEmbedder() inside a Web Worker, index with buildTurboQuantIndex() and search with searchTurboQuantIndex().
    Embed the documents once ahead of time, in Node.js or in the browser, and keep the index as a static file or in memory; a search then embeds only the query.
    The first use downloads the pinned Winzling model (39.8 MB) and caches it in the browser, so start it when the user first searches.
    ⏎ send

    Embed, index and search anywhere, for free

    Running an embedding model, building a vector index and searching it now works wherever Node.js or a browser runs: no GPU, no server, no per-query API bill.

    MIT licensed · Winzling model and Belebele data credited below

    Bench questions