Which Sally AI model is right for your meetings?
Sally transcribes with five AI models, and each one is built for a different job. Two write down every "um" because the exact wording matters. One is the everyday model, without the fillers and tuned for German-language meetings. Two are made for the hard recordings, where a lot of people talk at once and the audio is far from clean.
This page compares them side by side, so you can tell which one fits the way your team records. Which of them you can pick depends on your plan. Where you actually set the model is described under transcription model.
Sally's AI models are never trained on customer data. Customer audio, transcripts, and summaries are not used for training and are not accessible to us as a provider.
Quick navigation
- 1. The five models at a glance
- 2. The models in detail
- 3. What your plan includes
- 4. Which model to choose
- 5. How the models are trained
1. The five models at a glance
| Model | Filler words and vocal sounds | Speaker separation | Accuracy | Speed |
|---|---|---|---|---|
| Jade v1 | Yes | Strong | Very high | Standard |
| Jade v2 | Yes | Very strong | Very high | Standard |
| Opal v1 | No | Good | High | Standard |
| Topaz v1 | No | Strong | High | Fast |
| Topaz v2 | No | Very strong | Very high | Fast |
All five run entirely on the Sally platform, so no part of a recording is handed to a third party for transcription.
Next to the five there is Automatic, marked Recommended, which is what a new account starts with. Sally then looks at each recording and picks the model that fits it, so you do not have to think about it every time.
2. The models in detail
2.1. Jade v1
Our most detailed model, and the one that stays closest to the spoken word.
- Transcribes verbatim: filler words such as "uh" and "um" and non-verbal cues such as "(laughs)"
- Ratings: accuracy very high, speaker separation strong, speed standard
- First choice for: research interviews, advisory conversations and anything that has to be quotable word for word later on
That is the point of it rather than a side effect. In a qualitative interview or a sensitive conversation a hesitation is information, and Jade v1 keeps it on the record.
2.2. Jade v2
The next generation of Jade v1, built for rounds with many voices.
- Transcribes verbatim: like Jade v1, filler words and non-verbal cues included
- Ratings: accuracy very high, speaker separation very strong, speed standard
- First choice for: focus groups, advisory boards and group discussions
The step forward is in the speaker separation, which is the part that counts in a verbatim record: every sentence stays with the person who said it even when several people talk over each other. You can quote straight from the transcript instead of listening through the recording again.
2.3. Opal v1
The all-round model for everyday use, tuned for German.
- Transcribes clean: without filler words and non-verbal sounds, so the transcript reads smoothly
- Ratings: accuracy high, speaker separation good, speed standard
- First choice for: German-language meetings and day-to-day business
If your meetings run in German and you care about how the German reads, this is the one. Sally’s own summary of it: "Optimized for German-language online meetings (DACH), best German output".
2.4. Topaz v1
The proven generation of the Topaz line, broad and fast.
- Transcribes clean: without filler words and non-verbal sounds
- Ratings: accuracy high, speaker separation strong, speed fast
- First choice for: international rounds and conversations that switch language mid-sentence
Its particular strength is breadth: Topaz v1 holds up across many languages and audio conditions. Topaz v2 builds on it with the newest engine version.
2.5. Topaz v2
Our most precise model, and the one for the hard cases.
- Transcribes clean: without filler words and non-verbal sounds, like Opal v1 and Topaz v1
- Ratings: accuracy very high, speaker separation very strong, speed fast
- First choice for: meetings and workshops with several speakers in rooms that are not quiet
It stays robust with loud recordings and many participants, and it keeps the voices apart while it does. That is what earns it the top ratings for both accuracy and speaker separation.
3. What your plan includes
Not every plan offers every model. The more capable models come with the higher plans:
| Plan | Models you can pick |
|---|---|
| Starter | Jade v1, Opal v1, Topaz v1 |
| Pro | Jade v1, Jade v2, Opal v1, Topaz v1 |
| Enterprise | Jade v1, Jade v2, Opal v1, Topaz v1, Topaz v2 |
So every plan can transcribe verbatim with Jade v1, get a clean German everyday transcript with Opal v1 and cover mixed-language rounds with Topaz v1. Pro adds Jade v2 on top of everything Starter has. Topaz v2 is reserved for Enterprise.
Automatic is available on every plan and picks from the models your plan actually includes.
4. Which model to choose
The short version:
- You are not sure, or your meetings differ a lot → Automatic
- The exact wording, tone and hesitations matter → Jade v1
- Verbatim, but with several people in the room → Jade v2
- You want a clean, readable transcript of everyday meetings, in German above all → Opal v1
- Your recordings span several languages, sometimes within one sentence → Topaz v1
- Many speakers, noisy rooms, audio you cannot control → Topaz v2
4.1. Our general recommendation
Leave the default on Automatic and let Sally decide per recording. Meetings are rarely all alike, and the model that suits a quiet one-on-one is not the one that suits a workshop with eight people in a room with a hard ceiling.
Set a fixed model when your recordings genuinely share one characteristic and you want that characteristic served every time: an interview practice runs on Jade v1, a German-speaking sales team runs on Opal v1, a company whose meetings are large and loud runs on Topaz v2.
4.2. Choosing for one recording only
None of this locks you in. A single appointment can use a different model than the organization default, and a transcript that already exists can be written again with another model. Both are described under transcription model and optimize transcription.
5. How the models are trained
Protecting your data is a top priority at Sally. We never train our AI models using customer data.
Neither audio recordings nor transcripts or summaries from customer systems are used for model training, and they are not technically accessible to us as a provider.
5.1. No use of customer data
In concrete terms, this means:
- Customer data is not stored, analyzed, or reused.
- No training, fine-tuning, or prompt learning is performed using customer data.
- Meeting content remains the exclusive property of the customer.
- Processing is isolated, purpose-bound, and compliant with applicable data protection and security standards.
5.2. Training with proprietary, controlled datasets
Our AI models are trained exclusively on internal, anonymized, and legally compliant datasets, including:
- More than 120,000 hours of self-generated and licensed audio material.
- Over 25 million annotated sentences for transcription, speaker separation, and contextual understanding.
- Thousands of simulated meeting scenarios with varying audio quality, number of speakers, and domain-specific language.
- Controlled datasets covering accents, dialects, and industry-specific terminology.
These datasets are continuously expanded, reviewed, and quality-controlled, without using real customer conversations.
5.3. What this means for you
- Maximum data security with no hidden secondary use.
- Reproducible and explainable model quality.
- A clear separation between product usage and model training.
- Suitable for use in sensitive, regulated, or confidential environments.