Skip to content
Summo

Compare

Summo against Granola, Humla, Otter.ai and Fireflies — including the rows we do not win. With the measurements: 2.4% Vietnamese WER, 9× faster than realtime, and the command to re-run each one yourself.

Against the apps of the same kind

Not every comparison table is fair, so this one lists only what is checkable from their own pages — including the rows we do not win.

SummoGranolaHumlaOtter.aiFireflies
Audio leaves your machineNeverYes — sent to Deepgram/AssemblyNo, with a local modelYesYes
Works with no networkYesNoYesNoNo
A bot joins the callNot neededNot neededNot neededYesYes
Monthly minute capNoneYes, on the free tierNoneYesYes
Where your data livesMarkdown filesTheir US AWS VPCSQLite on your MacTheir serversTheir servers
Open sourceAGPL-3.0NoMITNoNo
Runs on LinuxYesNoNoNoNo
PriceFree$14/userFree · $5 for cloud$16.99/user$18/user

The two middle columns are apps of the same kind as Summo rather than cloud services — added because they are what you are actually choosing between, and because there are rows we do not win. Granola needs no bot either, and Humla is open source too. Sources: Granola's pricing and security pages ("best-in-class transcription providers (like Deepgram and Assembly)", "Notes are stored in our US-hosted AWS Virtual Private Cloud"), Humla's README and LICENSE. Otter and Fireflies prices are their lowest paid tier at the time of writing, billed monthly.

Models, like ollama pull

There is no "basic" and "premium" tier. Each model is a standalone manifest stating its size, the RAM it needs, its measured speed per CPU class, the languages it covers and its licence. Summo measures your machine and suggests one — you can pick a different one.

Vietnamese, English, and 97 other languages

A Vietnamese-specific transducer for the best speed and accuracy; Whisper for everything else and for sentences that mix Vietnamese and English. Switching models mid-recording does not lose the transcript.

An open registry, with nothing locked to us

Manifests are static JSON in a public repository. If Summo's CDN disappears the app falls back to GitHub and then to the original on HuggingFace. You can host the registry yourself.

Techainer/summo-registry

Measured, not promised

Every figure below has a command that reproduces it. Where it was measured is in the caption; you can run it yourself.

summo — recording · 12:04

rtf 0.11 · 9× realtime · ram 312 MB · gipformer-65m · offline

2.4%
Vietnamese WER (Fleurs VI)
summo-bench asr
9×
faster than realtime
summo-bench asr
0.94
voice-activity F1
summo-bench vad
30ms
searching 1,000 meetings
summo-bench vault
Recognition — gipformer-65M, 4 threads
DatasetWERCERRTFNote
Fleurs VI (15 clips, 146s)2.4%1.7%0.024read speech
A real meeting recording (34s)——0.107includes re-decoding for live text

That 0.107 includes re-decoding the sentence in progress about ten times so the text follows the speaker. Decoding once costs 0.012 — nine tenths of the budget buys words appearing as they are said, and there is still 9× headroom.

Storage — is a database needed?
MeetingsOn diskSearch (8 threads)List the vaultIndex size
1007 MB4 ms1 ms8 MB
1,00065 MB30 ms5 ms80 MB
5,000327 MB140 ms26 ms402 MB

The answer is no. Five thousand meetings is 327 MB of Markdown, searched in 140 ms — inside the 200 ms where a person starts to notice waiting.

No praise yet, only evidence

Summo has just been opened up and has no users to quote. Inventing a few kind words is easy, and it is the one thing a privacy tool must not do: you cannot open with a person who does not exist and then tell somebody you do not lie about where their words go. So here is what can be checked instead.

“BẠN SẼ DỄ DÀNG HỌC TIẾNG BỒ ĐÀO NHA”
What the end-to-end test read back out of the recording it had just made — verbatim, uncorrected.

A Vietnamese sentence, through a real microphone

Not a sample file: the test machine speaks out loud, records itself, and reads it back. whisper-tiny, on a CI runner, with no network.

pnpm -C apps/web e2e:full-flow

1,340 Rust tests and 384 interface tests

Plus 22 suites driven in a real browser against a real daemon and a real vault, on every push. None of them merely check that the code compiles.

cargo test --workspace && pnpm -C apps/web e2e

The packaged app is started before anybody is offered it

The AppImage is launched under Xvfb in CI: the daemon comes up, the interface is served, and the process is still alive five seconds later. The Android build is installed on an emulator and opened.

scripts/smoke-desktop.sh · scripts/android-smoke.sh

Open source, AGPL-3.0

The app, the engine and the builds are all on GitHub. Anything this page claims, you can read back in the code.

github.com/Techainer/summo-app

If you try it, say something

It worked, it broke, it was slow on this machine — all of it is useful. Real words will take this space.