← Back to VoltType
How the AI in VoltType works
Last updated: 27 September 2026 · Privacy policy · Terms
If your supervisor, your data protection officer or a procurement form has asked what the AI in this tool is, this page is the answer. Every model is named, with who made it, where it runs and what it receives. The numbers are measured, and the things that do not work are on the page too.
The six AI components. There is no seventh.
| What it does | Where it runs | Model | Made by | What it receives |
| Speech to text, on your machine | your own PC or phone, on the processor | Parakeet-TDT 0.6B v3 | NVIDIA, open weights | audio — never leaves the device |
| Speech to text, over the internet | Groq, in the United States | Whisper large v3 turbo (and large v3 for translation and some languages) | OpenAI, open weights, run by Groq | the audio, in transit only |
| The writer on your machine — AI Notes, meeting notes, spoken commands | your own PC, on the processor | Qwen2.5-1.5B-Instruct | Alibaba, Apache-2.0 | the text — never leaves the device |
| The writer over the internet — the same jobs | Groq, in the United States | gpt-oss-20b | OpenAI, open weights, run by Groq | the text, in transit only |
| Telling speakers apart in a meeting | your own machine | 3D-Speaker CAM++ embeddings, 28 MB | Alibaba DAMO, Apache-2.0 | audio — never leaves the device |
| Recognising a voice you have named, in a later meeting | your own machine (Windows) | the same CAM++ embeddings, saved | — | one saved voice fingerprint per name you typed |
Everything else in the product is ordinary code: the punctuation and number rules, the spoken-command tables, the silence gate, the plan counters. Those are tables written by hand, not models.
Three things we do not do, and one we will not build
- We do not detect emotion. VoltType does not judge mood, stress, health, attitude or character from a voice or from words. Checked across all 22 of the writer's instructions and the whole source tree. Inferring emotion in a workplace is one of the AI Act's outright prohibitions, and it is a permanent design ban here, not an oversight we might reverse.
- We do not train on anything you say. There is no training or fine-tuning pipeline in the product. All five models above are used exactly as their makers published them.
- We do not make decisions about people. Nothing here scores, ranks, screens or profiles anybody. The AI turns speech into text and writes drafts from it.
- Your words are not stored on our servers. No audio. No transcripts. Our usage table holds counts — seconds, words, which language, which model — and has no column that could hold content.
Voices you name, which is the part worth reading twice
If you type a person’s name next to a speaker in a meeting, the app can keep a voice fingerprint so the next meeting recognises them. This is the most legally significant thing in the product, so here it is in full:
- A fingerprint is a list of numbers. It contains no audio and cannot be turned back into a voice or into words.
- It is stored only on your own machine, in one file in your own folder. We never receive it. We never see it, hold it or back it up.
- Nothing is saved unless you type a name. Without a name a speaker stays “Speaker 1” and no fingerprint is kept.
- It refuses to guess: a name is used only on a clear match that also clearly beats the runner-up. Otherwise the speaker stays unnamed.
- Under EU law a voice fingerprint is personal data belonging to the person whose voice it is, and a special category at that. The person is usually not you — it is someone else in your meeting. The app therefore asks you to confirm they agreed, and you can delete any saved voice with one tap.
Recording people
VoltType records only when you press record. Everyone in the room must know and agree. In some countries — Germany and France among them — recording a conversation without everyone’s agreement is a criminal offence. The apps ask you to confirm this before your first recording, and say so again on the meetings screen. See section 9a of the Terms.
Where anything leaves the EU
One destination, and only while you are working over the internet: Groq, in the United States, which runs the two cloud models above. Audio and text pass through in transit and are not retained. Your account lives in the EU (Frankfurt). Payment goes to Stripe and account e-mail to Resend. Every one of them is named in the privacy policy, with the transfer safeguards.
With the on-device packs on Windows and Android, none of that happens at all: the audio and the text stay on the machine.
The measurements, including the ones that do not flatter us
- 8.2× faster than real time, on a processor with no graphics card. A 60-second dictation becomes text in about 7 seconds. Measured 26 September 2026 over 216 recordings, 13,031 reference words, 17 languages, 2.9 hours of audio, through the app’s own pipeline rather than a library on its own. By language the range was 4.4× to 12.5×. Measured on one PC with 20 cores using 8 engine threads; a slower machine will be slower.
- Three of the 25 on-device languages do not work on the device. Measured in the same run: Latvian 19.5%, Maltese 8.4%, Estonian 2.6% word accuracy. The voice pack falls back to English phonetics and cannot form their letters, so what it writes is not the language. Over the internet all three are good. They are marked wherever you meet them — on our pricing page and in the app’s language picker. We kept them in the count of 25 and told you instead of quietly dropping them.
- “Nothing left the machine” is a measurement here, not a slogan. Over three days in offline mode the app made 1,815 requests internally and sent 0 bytes out. The app keeps its own data log, so you can check the same thing yourself on your own machine rather than taking our word for it.
- What we have not measured, we do not claim. Accuracy percentages understate the strong languages, because a transcript that writes the same meaning differently is scored as an error. We publish the method with the number so you can judge it.
Where we stand on the EU AI Act
- Not high-risk. VoltType is a dictation and writing tool. It does not take or support decisions about people, which is what Annex III turns on.
- Not a prohibited practice. No emotion inference, no biometric categorisation, no social scoring, no manipulation.
- Article 50(2) applies to us and we have started early. Text a machine wrote is marked as machine-written where you see it and in anything you share out of the app. The legal deadline for products already on the market is 2 December 2026; the Android app shipped the marking on 27 September 2026 and the other apps are following.
- Article 4, AI literacy. This page is part of how we meet it. If you need something more formal for a procurement file, write to us and ask.
- Nothing to register. There is no EU database entry required for a system like this, and we hold no security certifications — we have had no external audit, and we would rather say so than imply one.
If a form asks you for a one-line answer
“An EU-hosted account, one American AI processor used only while online, and everything else on the user’s own machine. No training on user data, no emotion detection, and no audio or transcripts stored by the vendor.”
Every statement on this page was read out of the running system on 26–27 September 2026 — the source, the live database schema and the app’s own model files — not out of a marketing plan. If something here ever contradicts what the product does, the product is what counts and we want to know:
[email protected].
Contact
Questions about the AI, or a procurement form that needs filling in? Email [email protected].