I know this sounds like a privacy pitch. It is. But it's also just something true that most people have never bothered to think about. Every AI tool you've been using has been sending your documents to a server. The model reads them. A system logs them. Some of it trains future models. Keel runs entirely on your GPU, inside your browser tab. Nothing leaves. Not the question. Not the answer. Not the filename.
I'm not trying to scare you. Well, maybe a little. But this isn't paranoia. It's literally just what the terms say. Most people just never scroll past the headline.
The biotech founder at 11pm asking an AI to structure the patent filing she's three years from submitting. The M&A lawyer typing deal questions he can't ask his colleagues. The therapist drafting session notes with a patient's name in the document. Every one of those queries hit a server. The model read it. A system logged it. Some of it will train future models.
"But the companies say they don't—" They do. Open the terms. Scroll past paragraph eight.
The AI was never the product. Your information was. That's the business model. That's how the GPUs get paid for. It's not malicious. It's just the deal you agreed to without reading it.
Keel is a different deal. It's free. Ymy GPU runs the model. Your disk holds the documents. Your browser contains the conversation. When you close the tab, it's gone the way a thought disappears when you stop thinking it.
Not archived. Not retained. Not flagged for review by someone you've never met.
most RAG tools do one pass and hope. keel breaks your question into sub-queries, runs them in parallel, then scores every candidate through a cross-encoder before writing a word. the lawyer who asks "what are the indemnification obligations?" gets clause 14.3(b). not a paragraph that gestures near it. every sentence traces to a source.
the document, the query, the answer. none of it was ever on a wire. you don't have to trust a privacy policy. there's nothing for it to cover.
keel reads your GPU at startup, loads the largest model that fits your VRAM, and begins. 7B on a base M3 or RTX 4060. up to 70B on an M3 Max or RTX 4090. not throttled by someone else's queue. your GPU, full speed, answers the moment you hit enter. it tells you exactly which model you're running.
the founder who uploads three years of board decks and asks "when did the company first discuss a Series B?" gets a precise answer. not because it's in the cloud. because a real vector database with BM25 and semantic hybrid search ran the query locally. restart your browser. it's still there. always was.
PDF, Word, Excel, Markdown, plain text, code. four browser workers parse in parallel. chunks overlap at boundaries so nothing gets cut. a 50-page document is ready in under 4 seconds. after that it lives on your disk until you delete it. nobody else ever sees it.
keel reads your GPU at startup and loads the largest model that fits your VRAM. on a base M3 or RTX 4060: 7B, quantized. on an M3 Max or RTX 4090: up to 70B. as the hardware gets bigger, the model gets bigger with it. different hardware. same guarantee: nothing leaves the machine.
flip one toggle and keel stops being a search box. it splits your question into parts, investigates each against your documents, then reads its own findings and asks what's missing. the gaps become the next round. it repeats until a critic is satisfied or there's nothing left to find.
a cloud tool meters this. every extra pass is another line on the bill. keel runs all of it on your GPU, so there is no meter. the longest investigation and the shortest one cost the same. nothing.
a single merger-agreement question returned five findings and a fully cited report. you watch every step happen live.
there was. it just took a while to build. no waitlist. no signup. open your browser and use it.