Skip to main content

How to set the transcription model in Sally AI

The transcription model is the part of Sally that turns speech into text. Sally ships with several of them, and they differ in how precisely they transcribe, how well they tell voices apart, and whether they keep filler words. Under Transcription you pick the one Sally uses by default for new recordings.

The setting applies to the whole organization, so it is where an owner or admin decides what everyone starts from. Nobody has to think about models for a single meeting: the default carries every new recording, and a single appointment can still use a different model when it needs one.

For owners and admins

Only owners and admins can change the default transcription model. The choice applies to the entire organization, not just to your own account.

Quick navigation

  1. Who this setting is for
  2. Open the transcription settings
  3. The available models
  4. A different model for a single appointment
  5. Common questions

1. Who this setting is for

Only owners and admins can change the default transcription model, and that makes sense: the setting does not affect one person, it affects every new recording in the organization.

The benefit is that quality stops being a matter of chance. Instead of hoping each user picks a sensible model, you set one default that fits the way your company records, and everyone inherits it. If most of your meetings are German, you can say so once. If your recordings are noisy and full of speakers, you can pick the model built for that.


2. Open the transcription settings

To get there:

  1. Open Settings at the bottom left of the sidebar.
Sally sidebar with the Settings entry at the bottom left marked
Figure 1: Settings sits at the bottom left of the sidebar
  1. Under Configuration, pick Meeting Assistant, then the Transcription section.
Transcription section in the Meeting Assistant with the selectable transcription models and their ratings
Figure 2: The Transcription section with the selectable models
What the section controls

The info icon next to the title sums it up in one sentence: "Choose which transcription model Sally uses by default for new recordings."

Nothing takes effect until you click Save. The line under the heading always shows which model is active right now, and the Defined organization-wide switch marks the setting as binding for everyone.


3. The available models

You get one option per model, and each one carries a short description plus three ratings: Accuracy, Speakers and Filler words. Accuracy is how closely the text follows what was said, Speakers is how reliably Sally tells the voices apart, and Filler words says whether things like "um" and "you know" end up in the transcript.

3.1. Automatic

Automatic is marked Recommended and is the option Sally starts with. Rather than fixing one model for everything, Sally looks at each recording and picks the model that fits it, which is why the ratings column reads Sally decides per recording instead of showing fixed values.

This is the right choice for most organizations. Your meetings are rarely all alike, and a model that suits a quiet one-on-one is not the one that suits a noisy workshop with eight people. Pick a fixed model when you know your recordings share one characteristic and you want that characteristic served every time.

3.2. The five models in detail

ModelWhat it is built forAccuracySpeakersFiller words
Jade v1Verbatim and nuanced, for interviews and sensitive conversationsVery highStrongYes
Jade v2Like Jade v1, with clearly better speaker separationVery highVery strongYes
Opal v1Optimized for German-language online meetings (DACH), with the best German outputHighGoodNo
Topaz v1Robust across many languages and audio conditions, including mixed-language recordingsHighStrongNo
Topaz v2Highest accuracy, for many speakers and noisy recordingsVery highVery strongNo

Jade v1 is the only model that keeps filler words and vocal sounds. That is the point of it: in an interview or a sensitive conversation, a hesitation carries meaning, and a transcript that quietly tidies it away loses something. For a status meeting the same behavior just makes the text longer.

Jade v2 is the next generation of it. It stays just as faithful to the wording, but it separates the voices considerably better. That makes the verbatim record usable for conversations with several people as well.

Opal v1 is the German specialist. If your meetings run in German and you care about how the German reads, this is the model tuned for it.

Topaz v1 is the all-rounder. It holds up across many languages and it copes when a conversation switches language mid-sentence, which the other models handle less gracefully.

Topaz v2 is built for the hard cases: a lot of speakers, a noisy room, audio that is not clean. It carries the highest ratings for accuracy and speaker separation.

Not every plan offers every model

Which models you can pick depends on your plan. Starter covers Jade v1, Opal v1 and Topaz v1. Pro adds Jade v2, and Topaz v2 is reserved for Enterprise. Automatic works on every plan and picks from the models your plan includes.

The settings show three of the ratings. The full side-by-side comparison, which also rates how fast each model works, is under AI models.


4. A different model for a single appointment

The organization default is a starting point, not a cage. Sally says so directly under the model list: "This is the default for new recordings. For a single appointment you can pick a different transcription model in the appointment settings."

So a single appointment can run on its own model without anybody touching the organization setting. The appointment settings offer the same comparison, and Automatic there means the appointment falls back to what the series or the organization defines.

For a recording that already exists, the model is not fixed either. You can have Sally write the transcript again with a different model, which is described under optimize transcription.


5. Common questions

5.1. Who can change it

Owners and admins, nobody else. For what each role is allowed to do, see user role management.

5.2. How far the setting reaches

Across the whole organization, and only for recordings made from then on. Changing the default does not touch transcripts that already exist. The Defined organization-wide switch is what marks the choice as binding for everyone.

5.3. Recordings that already exist

They keep the model they were transcribed with. If you want a different result for one of them, have Sally transcribe it again, as described under optimize transcription.

5.4. Why a model is missing

Because your plan does not include it. Starter offers Jade v1, Opal v1 and Topaz v1. Pro adds Jade v2, and Topaz v2 comes with Enterprise. The full comparison is under AI models.