Skip to content

Model scoring

Models are ranked on production data, not on their own benchmarks

Five vendors, five leaderboards, all of them flattering. We believe none of them. The only thing that counts is what a vendor actually delivered to a user on our line. One thing up front: no vendor credentials are configured on the public environment yet, so this scoreboard is empty. What follows is how the score is computed, and why the numbers will be worth trusting on the day there are any.

Why not the vendor benchmarks

They are not sitting our exam

A vendor benchmark is their own leaderboard: they choose the samples, the subject matter and the marking scheme. That exam looks nothing like a vertical drama shot — one that carries a reference image, runs two seconds, and needs readable Chinese on a shop sign. Those questions are not on their paper. More to the point, a leaderboard never answers the question the line actually has: a cheap model that fails one time in three costs more, once you count the retries, the waiting and the burnt quota, than the expensive one you were avoiding. No vendor table will ever tell you that.

  • There is one source of raw data: real renders. Whether it succeeded, how long it took, how many times that shot was redone afterwards, whether anyone marked it down.
  • Every call writes a usage record, including the failed ones (zero units, so nothing is billed). Skip that and the denominator only contains successes, and everybody scores 100%.
  • Only the last 30 days count. Everything inside the window counts equally, everything outside is simply absent. A vendor that was bad last month and fixed it this month catches up within a month.

Four signals

How the weight is split

Reliability weighs the most, because a model that is fast and good but fails one time in three is a net cost on the line: you wait, you retry, and you burn quota doing it. Speed weighs the least, because slow is only uncomfortable. Failing is what stops a film shipping.

Reliability45%

Measured from

Whether each call succeeded, and whether it fell back

A success rate is an objective fact; there is no dressing it up. A fallback counts as half a failure: it delivered something, but not the thing you paid for.

Rework25%

Measured from

How many times the same shot got rendered

When you redo a shot you are telling us that one was not good enough. The signal costs you nothing, because that redo was going to happen anyway.

Satisfaction20%

Measured from

Your thumbs up and down on individual shots

Explicit feedback: low volume, clear direction. A thumbs-down picks a fixed reason rather than free text, because “character drifted” and “movement looks wrong” point at completely different model weaknesses, and only fixed reasons aggregate into anything meaningful.

Speed10%

Measured from

Wait per second of finished footage

The lightest weight, but it genuinely affects the experience, so it stays on the table. Waiting is normalised per second of footage so that long shots are not penalised for being long.

The one that matters most

Rework: costs you nothing, and cannot lie

Of the four signals, rework is the one we lean on hardest, for one reason: it is behaviour, not opinion. Ask someone to rate a shot and most people skip it. But nobody leaves a shot they are unhappy with in the film. The moment you press “redo this shot” you have given the most accurate verdict available, and it cost you nothing at all.

  • Only shots that actually rendered count. Retries after a failure do not count as rework, because that is the system's problem, not you saying the result was bad.
  • Which vendor a thumbs-down points at is read from the output's own record, never supplied by the caller. Otherwise anyone could downvote a competitor.
  • Scores are separable per team. The platform-wide view has the most samples and drives the routing default; your own team's view reflects your own subject matter — a thriller and an ad spot do not ask the same things of a model.

Five rules

A score that can be gamed is worse than no score

These five are not implementation details. They are the precondition for trusting any of this. Each one blocks a way of making the numbers look better while making the choice worse.

01

Too few samples means no score

Under 8 calls we report the raw data only: no composite, and no influence on routing. A success rate over three calls is noise, not signal. Better to print “not enough data” than to let it decide who renders your film.

02

Missing signals are excluded, never counted as zero

“Nobody has rated it yet” is not “rated badly”, and “we did not record the latency” is not “slow”. A missing signal is dropped and its weight is shared out proportionally among the survivors. With only reliability and rework left, the 45-to-25 ratio between them is preserved exactly.

03

New models start neutral

A vendor with no data starts at a neutral 0.6, not at zero. Score them zero and anything newly connected sorts last forever, never gets a sample, and therefore never gets a score. The scoring system becomes a self-fulfilling prophecy.

04

A fallback counts as half a failure

Dropping to the built-in preview cannot be recorded as a success. It did deliver something, but not the thing you paid for. Counting it as a win would make a vendor that never errors but always falls back look flawless.

05

The score orders candidates; it never removes them

A low-scoring vendor can still be picked, especially when the others are down. Our floor is that a film ships, not that only the best vendor is ever used. A system that gives up because no candidate is good enough is worth nothing to you.

Standing benchmark

A new model should not have to wait for work to get a score

Scores come entirely from real production, which is unbiased but starts new models cold: a vendor needs a few jobs before it has a score, and without a score it does not easily get jobs. So we run a set of standard cases against each vendor deliberately, and that vendor has a baseline the same day. Lab data and real production stay separable in the report, because standard cases are not your actual subject matter.

Six video cases

01
Static plate

Whether the picture holds still and is free of obvious artefacts. This is the baseline — fail it and nothing else matters.

02
Human movement

Whether movement is continuous and limbs stay attached. This is where models give themselves away fastest.

03
Camera move

Whether it actually performed the requested move, or just handed back a still with some shake on it.

04
Text in frame

Whether Chinese glyphs come out correctly. Many models are poor at this, and vertical drama is full of signs and doorplates.

05
Character consistency

Whether it is the same person as the reference image. This one decides outright whether serialised work is possible.

06
Fast short shot

Whether a short shot gets stretched or padded with extra action. Ads and music videos live on this.

There is exactly one bar for adding a case: can it separate the models? “Generate a landscape” is something everybody can do, so running it is just burning money. “A shop sign with two Chinese characters on it” tells you at a glance who can and who cannot. A run costs real money — cases × vendors × seconds each — so it prints an estimate and asks before it starts.

The score is for the router. The film is for you

All of this serves one purpose: sending each shot to whoever is most likely to get it right first time, right now. You never need to know who that was. But if you want to, the console says so.