How it worksPrivacyExperts RoutinesThe name FeedbackGet the app
Free · no account · nothing leaves your phone
ternii app iconternii
0tokens a second.
On an iPhone. In the app.iPhone 17 Pro · Metal · in 195 MB peak · measured

Private AI, light enough to fly.

An assistant that runs entirely on your phone. Ask it anything — on a plane, on a train, off-grid — and nothing you say ever leaves the device.

0 bytes to the cloud by default 61–74 tokens/s on a Galaxy S25+ · in the app answers begin in 190 ms a 975 MB model that costs 195 MB to run 0+ on-device experts · growing runs on a phone coming up to ten years old · iPhone X, 3 GB RAM English at launch · more languages later

Store links are placeholders while we finish review. Android first; iPhone follows.

terniioffline
Ask anything — it stays on your phone
▶ live demo · running on-device, in airplane mode
Why it can live on your phone

Three states. Radically lighter.

Ordinary AI does its sums in fine-grained numbers — expensive to store, expensive to move, expensive in energy. ternii's models think in just three: +1 0 −1. That makes the whole assistant small enough, and frugal enough, to run offline on a phone — no data centre, no bill, no waiting on a network.

🪶

Small by design

A whole assistant in about a gigabyte. It loads into your phone's memory and answers there — nothing to upload, nothing to download per question.

Fast where it counts

Three-state maths moves far less data per word. On an iPhone 17 Pro, ternii's own model writes at up to 118 tokens per second in the app, and the first token arrives in 190 ms. On a Galaxy S25+, 61–74 tokens per second. Measured on the phones, in the shipping app — not projected.

iPhone 17 Pro typical 108–118 tok/s (bench 112); time to first token measured end-to-end from the model call to the first token emitted. Short runs on a cool phone; ±8 % run to run, so we quote ranges.

📴

Works in airplane mode

The model is on the device. Turn off Wi-Fi and mobile data and it still answers — on a train, on a plane, off-grid.

Same phone. Same app. Only the model changed.

In our app, on an iPhone 17 Pro, the difference is not subtle.

One phone, one app build, one prompt, only the model changed. Ours is ternary from training and runs on the GPU; a 4-bit model is a compressed copy of a bigger one and has no GPU path in our app — so this is an as-it-ships comparison, and we say so.

118

tokens a second

Llama-3.2-3B in the same app: 12–18. BitNet 2.4B (ternary, same GPU path): 53–56.

Best in-app run; typical 108–118. Model to model on the same backend: 3.1× Llama.

190 ms

to the first token

Llama-3.2-3B in the same app: 6.3 s. Qwen3.5-4B: 8.3 s.

33× and 44× lower time to first token, measured end-to-end.

195 MB

peak memory

Llama-3.2-3B: 596 MB. Qwen3.5-4B: 600 MB. From a file half to a third the size.

The weights are memory-mapped, not loaded — they cost nothing the OS charges you for.

~10 yrs

a phone coming up to ten years old still runs it

iPhone X (2017): 12.6–16 tokens a second in 240 MB. On a 2022 Galaxy S22, a 4-billion-parameter model manages 0.22 tokens a second — it doesn't usably run. Ours: 17–20.

One 975 MB model file, same hash on every phone.

Not on this page: answer quality. ternii's model is behind cloud-class 4B models on knowledge benchmarks — faster, smaller, runs where they don't, not yet as accurate. Every number above is a short run on a cool phone; we quote ranges, and say when a number is a best run — 118 is the best in-app run; typical is 108–118.

Ternary from the first token

Nothing lost to compression.

Most phone AI is a big model squeezed down to fit — and it loses accuracy in the squeeze. ternii's model was trained in three states, so the model on your phone is the model we trained. There is no compression step, and nothing to lose in one.

🎯

The model you get is the model we trained

Every weight was −1, 0 or +1 from the first token of training. Nothing is rounded off afterwards to make it fit a phone — the 1 GB file on your phone gives the same answers as the 7 GB file it was packed from.

Measured: the model file on your phone scores the same as the model we trained — held-out perplexity within 0.25 % of the training checkpoint, and identical HellaSwag (40.4 % vs 40.4 % on the same 2,000 tasks) to the 7× larger uncompressed file. 2026-09-08, same texts, same scoring.

📄

Native ternary holds its accuracy at scale

Published: a natively-ternary 2-billion-parameter model trained on 4 trillion tokens lands within about one point of a comparable 16-bit model on a ten-task average (54.2 vs 55.2) — Microsoft's published result, not ours.

Microsoft, BitNet b1.58 2B4T technical report (2025), vs Qwen2.5-1.5B. That is the recipe ternii's model follows — ours is smaller, earlier in training, and English-only for now, and we say so.

How it works

A small model, given the right facts at the right time.

A tiny model doesn't know everything — so we don't ask it to. ternii's trick is augmentation over size: routing and retrieval put curated facts and live tools in front of the model exactly when a question needs them, and the model writes from those facts only.

You ask — typed or spoken
On-device speech-to-text, or just type. Nothing is sent anywhere.
Route — the app decides where the answer should come from
Maths? A fact? A live tool? A step-by-step problem? Each takes a different path — chosen in code, not left to the model to guess.
Ground — pull the exact facts that fit
The best-matching expert rows, a device connector, or a web lookup you approved — each answer shows its source.
Answer — written from those facts only
When it genuinely doesn't know, it says so and offers to search — instead of making something up.
🎯

Honest by construction

Exact arithmetic is computed, not guessed. Questions about your private data, your history, or the future are declined rather than invented. Every grounded answer is cited.

🧭

Pick your model

ternii's own model is the fast default. A larger open ternary model is one tap away when you'd trade some speed for depth — and you always see which model answered.

🔎

Search, on your terms

Every reply can reach out to the web for a cited summary — but only when you tap it. Offline is the default; online is opt-in, and only your query ever leaves.

Privacy & on-device

Your phone is the whole computer.

ternii is private by construction, not by promise. There's no account to create, no profile to build, and no server that sees your conversations, your notes, or your files. In airplane mode it works exactly the same.

🔒

Offline by default

Everything runs on the device. A single toggle turns web lookups on, with a clear prompt — and only the query leaves, never your chat or your data.

🙅

No account, no tracking

Download and run. No sign-up, no ads, no analytics profile. What's on your phone stays on your phone.

🗄️

Encrypted at rest

Your session, notes and settings are stored encrypted on the device. You can wipe them any time.

A panel in your pocket

Experts, downloaded only if you want them.

Knowledge lives in small, curated expert packs — one per subject. The app stays light and you add the packs you care about; each is 55–334 KB and installs in a tap. New packs appear over time, with no app update needed.

MathematicsPhysicsChemistryBiology Medicine & HealthComputer ScienceEngineeringHistory GeographyLawLegislationEconomics & Finance PsychologyPhilosophyCooking & FoodHow-To Writing & LanguageEarth & SpaceGeneral Knowledge

The expert room

With the larger model on, several experts can share one question — the model weaves their facts into a single answer and shows you which packs it drew on. A whole panel, deliberating on your device.

Simple agentic AI, on-device

Routines that do the legwork.

A routine gathers from the tools you've connected — calendar, weather, news, prices, your notes — and hands the model only the real data to summarise. No invented meetings, no made-up headlines: if it isn't there, it isn't in the briefing. One tap, a useful answer.

🌅

Morning briefing

Today's calendar, weather and headlines, in a few lines — before you're out the door.

🌙

Evening wind-down

Tomorrow's schedule and any loose ends from your inbox, tied off for the night.

Alarms & events

"Set an alarm for 6:30 tomorrow" — ternii asks once, then hands it to your clock or calendar. You confirm every action.

📰

News digest

Top headlines from your chosen feed, condensed to the bullets that matter.

📈

Market check

Your watchlist prices with a one-line read — no dashboards to open.

💡

Fact of the day

A genuine nugget pulled from one of your experts, explained simply.

one tern · two rails · three states
Why "ternii"

The whole idea, in one word.

Every part of the name is literally true of what's running on your phone.

tern

Ternary. The models think in three states — +1, 0, −1 — which is the reason a real assistant fits in a gigabyte and runs on a battery.

tern

The bird. The Arctic tern weighs about 100 grams and flies pole to pole every year, with no infrastructure at all. Small, light, goes anywhere. That's on-device AI.

ii

Two rails. Your phone's chip is binary, so each three-state value rides on a pair of binary signals — two rails. ternii is three states, carried on two rails, on the phone you already own.

Shape the panel

Request an expert.

Want a pack we don't have yet — beekeeping, tax, a language, your trade? Tell us. Because experts are just downloadable packs, we can add a new one for everyone without an app update.

Thanks — opening your email app to send it.

We're listening

Feedback, bugs & requests.

Found a bug? Want a new routine, or a feature? ternii gets better from what real people ask for. Tell us what happened or what you'd love to see.

🐛 Bug report⚡ Routine idea✨ Feature request💬 Just thoughts

Thanks — opening your email app to send it.

Free forever

Try ternary now.

A private AI in your pocket — and a first look at what three states can do. Download, and see what a phone can do on its own.