# engineering
Building Guardian Pocket: The Local Health Agent You Actually Own
This is Guardian Pocket. It's the world's first local health agent where YOU own the data end to end.
- It's 100% private. Data never leaves your phone.
- It's as good as expensive frontier agents at its tasks.
- It's blazing fast.
- To do it we fused model and harness into one integrated system.
Our goal with A-LIST is to make health easy. Several weeks ago we released a new version of A-LIST which is still AI powered but it requires no account, is free, and is totally private. None of your A-LIST data ever leaves your phone.
However, our AI health coach "Guardian" runs using a frontier AI model that costs money to serve and data has to leave your phone to get answers. We left Guardian out of the local version of A-LIST.
To bring it back we asked ourselves: Can we build a version of Guardian that does most of what people rely on it for today, but runs entirely on your phone?
The answer is yes. It's out now. It's called Guardian Pocket.
Here's all the engineering that went into it.
Base Model Selection
We surveyed hundreds of models and did full evaluations on dozens of them and found one clear winner. But first some notes about the losers:
- Anything over 12B parameters was too big on disk and took up too much memory.
- Anything under 6B parameters just didn't cut it on intelligence.
- MoE models performed much worse than dense models on our intelligent evals.
8B parameter dense models seemed like the narrow window where this would actually work, but we also needed it to be quantized fairly dramatically.
That's what led us to the winner on both speed, size and intelligence: Prism's Ternary Bonsai 8B.
Their quantization-aware training succeeded in bringing us the raw performance and intelligence we needed. Even just before releasing Guardian Pocket, we went back to retrain and eval other models on the full system and none came close to Bonsai.
Context Engineering — Solving Intelligence
When running an LLM on a phone every token counts. This improves response speed, but at this point we mainly cared about generation quality. Irrelevant tokens mean lower quality responses.
Our own experience over the past 18 months of building agents is that agentic search performs better than RAG in agentic systems so we started with a 2-query strategy.
- Context decision: The first query asks the model what it thinks it needs to respond to the user.
- Generation: The second query pulls in the extra context and generates the final response.
The first pass controls almost every part of the context that can be pulled in including:
- Agentic skills
- User files
- Previous conversation turns
- Extra user health data
As a result we've been able to keep every query under 2000 input tokens and get great responses from the generation step.
Generation Engineering — Solving Speed? (but actually solved hallucinations)
With this in place, we got great intelligence but turns could easily take 60 seconds. Not good.
Recently Jev has been making the rounds in the AI community with many projects like Kev applying those insights to get similar results. The intuition is that you can easily tune and modify an LLM to be a classifier instead of a generator. This involves getting into the guts of the model to get instant generation results after reading in the initial context.
To build your own "Jev at home" you need a "head" which is trained to take the raw outputs of an LLM pass before a token is generated and turn that into a classification decision. This works perfectly for our case because we need to decide what context to bring into our query.
Our first eval after training the head showed good results but not great results. So we also trained a Low-Rank Adapter (LoRA) for the model to help improve the classification results and achieve sufficient quality.
| Decision | Before: router or keyword gate | Trained head |
|---|---|---|
| Load lab results | 79% / 82% | 91% / 97% |
| Load Apple Health data | 75% / 92% | 100% / 92% |
| Offer logging | 67% / 97% | 88% / 97% |
| Offer reminders | 87% / 100% | 87% / 100% |
| Offer calculator | 46% / 40% | 52% / 87% |
| Decision | Linear head only | Head + aLoRA |
|---|---|---|
| Offer logging | 88% | 92% |
| Load Apple Health data | 80% | 89% |
| Offer calculator | 62% | 69% |
One happy side effect of this work is that we also solve hallucinations.
Hallucinations are mostly a thing of the past on the frontier but still happen on smaller models.
We were also able to classify for tasks the model couldn't handle and instruct the generation side appropriately. This was CRITICAL to delivering quality because queries that were outside of the model's ability to properly respond to would almost always result in a hallucination. Hallucinations were reduced dramatically.
| Request category | Before | With trained boundary decisions |
|---|---|---|
| Requests outside the phone's capabilities | 1.8 / 10 | 5.6 / 10 |
| Medicine questions | 4.5 / 10 | 6.9 / 10 |
| Emotional, general, and supported requests | — | Reported unchanged |
The whole effort gave us a very modest speedup but not the dramatic speedup we were hoping for. The problem is we still needed to read a huge amount of context into the model twice for the 2-query architecture with none of that work being reused.
Harness Engineering — Fusing the Harness and the Model to Solve Speed
Let's take a look at what we have so far:
- Query 1:
[global context block & user query] + [decision prompt]→ LoRA → Bonsai → Head →[decision] - Query 2:
[global context block & user query] + [extra context]→ Bonsai → decision
Where is most of the time going in query 1? Filling Bonsai with [global context block & user query].
Where is most of the time going in query 2? Filling Bonsai with [global context block & user query].
Yes that's the same thing twice. That's wasted work and mostly work that can be done before the user presses "send message".
So we built a harness that not only provides tools and an environment for the model to play in. Our harness actively manages the model's GPU memory and its layers, on the fly, in milliseconds.
Here's how it works:
- The user opens Guardian Pocket and the harness starts running Bonsai with the global context and system prompt. This is check-pointed for reuse later.
- The harness applies the head and aLoRA to the warmed model.
- User presses send which adds a few dozen tokens. A context decision is reached in < 1 second.
- The harness removes the head and aLoRA to prepare for generation.
- If there is no extra context needed generation begins on the second query < 500ms later.
- If extra context is needed, generation begins after a short delay.
The result is there are now cases where Guardian Pocket responds instantaneously. Guardian Pocket was responding so fast, we actually had to change our Guardian UX to deal with the speed.
Full Guardian Parity: What Comes Next
We have line of sight with existing techniques to get to full parity with the capabilities of our frontier Guardian. This will require us to specialize our local model for our specialized use cases. It will also require more harness work to bring the full toolset to our local Guardian.
In the coming months we'll begin:
- Mid-training models on high quality health knowledge.
- Post-training on mobile harness agentic uses.
- Use quantization-aware techniques to continue to shrink the model footprint and improve speed.
- Use jev-training techniques to build an optimal context classifier.
- Bring the full suite of Guardian tool calls.
A-LIST is a citadel for your health AND most sensitive health data. Your AI should be the Guardian of that citadel.