Microsoft models on UGESI

2 engines covering utility, audio. Every rate is published and shown again before you generate.

Florence 2

Detection, captions and OCR

1 cr / image

VibeVoice

Expressive long form speech

3 cr / 1k chars

What this lab is known for

Microsoft has 2 models here, covering utility, audio.

Rates run from 1 to 3 credits. That is the spread between a draft and a flagship.

Every one of them runs on the same balance as models from other labs. You are not locked into a vendor, and switching costs nothing.

Why the lab matters less than the job

Most people arrive here searching for a lab name. The job decides the result. A portrait needs different strengths than a product shot.

Every model below runs on the same balance. You switch between them in one tool. The rate is confirmed before each run.