Sarvaswa AI Labs

AI creative evaluation: computer vision and generative AI that review an ad before it ships.

Marketing teams at large consumer brands review thousands of pieces of creative a year by eye, and the judgement varies with who is looking. We built a platform that reviews a poster, banner or pack shot against a brand's own visual standards and returns a readiness verdict in seconds. Deep learning models trained on AWS SageMaker AI measure the creative. A generative model, run by our agent framework on Amazon Bedrock, interprets those measurements and explains what to change.

Built withAWS SageMaker AIAmazon S3Amazon BedrockCustom vision modelsSarvaswa agent framework

How it works

From an uploaded creative to a readiness verdict.

The design separates measuring from judging. Vision models produce facts about the canvas: what is present, how large, where, and in what contrast. A generative model then applies the brand's standards to those facts, decides each check, and writes the reasoning. Neither could do the job alone: a vision model cannot explain, and a language model cannot be trusted to measure.

  1. 01

    Upload

    A creative and its context

  2. 02

    Store

    Amazon S3, versioned with its verdict

  3. 03

    Measure

    Vision models on SageMaker AI

  4. 04

    Judge

    Bedrock agent applies the standards

  5. 05

    Explain

    Verdicts, reasoning, shareable report

  6. Every verdict carries the measurement behind it. The reasoning is generated; the numbers are not.

What changed

Manual creative review, and what replaced it.

Before the platform, each piece of creative waited for a reviewer, and each reviewer applied the standards a little differently. Feedback came back as opinion rather than evidence, and the same avoidable mistakes kept reaching production.

Manual creative review

  • A queue measured in days, with capacity limited by the reviewers available.
  • Verdicts that depend on the reviewer, the market and the time of day.
  • Feedback delivered as comments, with nothing a designer can measure against.
  • Recurring mistakes that nobody catches consistently across markets.

AI creative evaluation

  • A verdict in seconds, on every upload, with batch review for whole campaigns.
  • The same standards applied the same way every time, adjustable per product line and market.
  • Each check backed by a measurement a designer can act on.
  • A specific, plain-language recommendation for every check that fails.

Deep learning plus generative AI

Why the accuracy comes from combining two kinds of model.

Generative models are persuasive but cannot be relied on to measure. Trained vision models measure precisely but cannot explain. The platform gives each the job it is good at, and the hand-off between them is where the accuracy comes from.

What the vision models do

  • Detect the elements that matter to a brand: identity marks, packs, people, callouts and text.
  • Measure size relative to the canvas, position, contrast against surroundings, and how many elements compete for attention.
  • Read the copy and describe its shape: length, casing, and the kind of language used.
  • Return facts with bounding boxes, so every measurement can be checked by eye.

Trained and evaluated as SageMaker AI jobs against a large set of creatives labelled by the brand's own reviewers.

What the generative model does

  • Apply the brand's standards to the measurements, with thresholds that differ by product line and market.
  • Decide the checks that need judgement rather than a ruler, such as whether a layout reads clearly at a glance.
  • Write the reasoning for every verdict and a concrete recommendation for every fail.
  • Roll the checks up into an overall readiness score and a report the team can share.

Orchestrated by our agent framework on Amazon Bedrock, with each check run as a separate, evaluated step.

What is checked

The kinds of standard a creative is reviewed against.

The standards come from how the brand already briefs and judges creative, captured as a library of checks rather than a document nobody reads. Each check has a measurable basis, a threshold that can vary by product line and market, and a plain-language explanation of why it matters. The exact library is the client's; these are the families it covers.

Brand presence

Whether the identity marks and distinctive assets a shopper should recognise are present, large enough and prominent enough to register at a distance.

Measured by

Detection, relative size and contrast.

Layout and clarity

Whether the composition reads at a glance: how many elements compete, how they are arranged, whether the background separates them, and whether the hierarchy is obvious.

Measured by

Element counts, positions, contrast and structure.

Human attention

Whether people in the creative draw attention and direct it toward the product or the message rather than away from it.

Measured by

Face detection, prominence and gaze direction.

Message and call to action

Whether the copy is short enough to read, written in a way that is easy to scan, and asks the shopper to do something.

Measured by

Text extraction, length, casing and language pattern.

Each check returns a verdict with its measurement and reasoning. The verdicts roll up into an overall readiness score and a simple band, so a brand manager reads the position in a glance and the detail on demand, and downloads a report to share with the agency.

Creative going out every week?

Fifteen minutes to see how your standards could be measured.

Bring the guidelines your brand team already argues about and the creatives that slipped through. We will show you which standards a vision model can measure today, which need a generative model to judge, and what a proof of concept on your own creatives would return in two to four weeks.

Beyond creative

What else trained vision models can do for a business.

The same architecture, deep learning to measure and generative AI to judge and explain, applies wherever a person currently looks at images or video and makes a repeatable decision. These are the applications we build most often.

Shelf and planogram compliance

Is this store shelf laid out as agreed, and which facings are missing?

Field photos scored against the planned layout.

Packaging artwork checks

Does this artwork carry the right legal copy, barcode and claims for this market?

Pre-press review that catches errors before a print run.

Manufacturing defect detection

Which units on this line show a dent, a misprint or a missing seal?

Camera feeds scored in real time, with reasons a technician can act on.

Document and form extraction

What are the fields, tables and signatures in this scanned form, and are any missing?

Layout-aware OCR feeding a structured record.

Video and frame analysis

For how long is the brand visible in this ad, and where does attention go first?

Frame-level detection rolled up into a report.

Brand safety and moderation

Does this user-submitted image meet our brand and platform guidelines?

Classification with an explanation, not just a flag.

How we build one

How a vision build runs, from standards to a working platform.

A proof of concept on your own images typically lands in two to four weeks, and a production platform in eight to fourteen. The sequence is the same for any vision use case: agree what is being judged, train the models that measure it, add the model that judges and explains, then roll it out.

  1. 01

    Standards and data

    Weeks 1 to 2

    What happens

    We work with the people who review today to turn their standards into checks with a measurable basis, define what pass and fail look like for each, and collect labelled examples into Amazon S3.

    What you get

    Check library, labelled dataset, measurement plan.

  2. 02

    Train the vision models

    Weeks 3 to 6

    What happens

    Detection, text extraction and measurement models are trained and evaluated as SageMaker AI jobs against the reviewer labels, until they agree with the reviewers at least as often as the reviewers agree with each other.

    What you get

    Trained models, training pipeline, evaluation report against human verdicts.

  3. 03

    Agent and report

    Weeks 7 to 10

    What happens

    Our agent framework on Amazon Bedrock applies each check to the measurements, writes the reasoning and recommendations, and assembles the score and the report. Every check is an evaluated step.

    What you get

    Agent configuration, prompt registry, report template, per-check evals.

  4. 04

    Rollout

    Weeks 11 to 14

    What happens

    The web application goes to the brand teams: single and batch upload, context selection, history, and report download. Thresholds are tuned from the first weeks of real use.

    What you get

    Web application, batch processing, runbook, thirty-day support window.

What you get

What is in place when we hand over.

A working platform and everything needed to run and extend it, documented for the team that will operate it.

  • 01

    Trained vision models

    Detection, text extraction and measurement models with their SageMaker AI training pipeline, so they can be retrained as creative styles change.

  • 02

    The check library

    Every standard as a versioned definition with its thresholds per product line and market, editable without a code change.

  • 03

    The agent framework configuration

    The Bedrock agent, its prompt registry and the evaluated steps that turn measurements into verdicts and recommendations.

  • 04

    The web application

    Single and batch upload, context selection, history, per-check detail with annotated images, and report download.

  • 05

    An evaluation set

    Creatives with human reviewer verdicts, run before every model or prompt change, so accuracy is measured rather than assumed.

  • 06

    Runbook and support

    A runbook for adding a check, a product line or a market, plus a thirty-day post-launch support window.

AI creative evaluation, answered.

Software that reviews marketing creative, such as posters, banners and packaging, against a brand's visual standards and returns a readiness verdict with reasoning. The platform Sarvaswa built uses trained vision models to measure what is on the canvas and a generative model to judge each check and explain what to change. A review that took a brand team days now takes seconds.
Because each is unreliable at the other's job. A generative model asked whether a logo is prominent enough will answer confidently with no measurement behind it. A vision model can measure size, position and contrast precisely but cannot explain what to change. The platform uses vision models trained on SageMaker AI to produce the facts and a generative model on Amazon Bedrock to judge and explain them, so every verdict is both measured and reasoned.
Accuracy is measured against human reviewers. During the build we collect creatives with verdicts from the brand's own reviewers and evaluate every model and prompt change against that set, check by check. The target is agreement with the reviewers at or above the level the reviewers agree with each other. Measured checks are deterministic; judged checks carry reasoning the reviewer can verify against the annotated image.
Yes. The standards live in a versioned library rather than in code, with thresholds per product line and market. Adding a check means defining what a pass looks like, supplying labelled examples, and running the evaluation set. Different product lines and markets already score with different thresholds in the platform we built.
Static creative in common image formats: posters, banners, digital display, social assets and pack shots, singly or as a batch for a whole campaign. Video is reviewed frame by frame in a separate configuration, one of the applications described on this page.
Anywhere a person looks at images or video and makes a repeatable decision: shelf and planogram compliance from field photos, packaging artwork checks before print, defect detection on a production line, document and form extraction with layout-aware OCR, video frame analysis for brand presence, and brand safety moderation of user-submitted content. The architecture is the same: vision models measure, a generative model judges and explains.
A proof of concept on your own images typically lands in two to four weeks: the standards agreed, a first set of models trained, and verdicts on real examples. A production platform with a web application, batch processing, per-context thresholds and a runbook typically takes eight to fourteen weeks.
AWS, under the customer's own account: creatives and reports in Amazon S3, vision model training and inference on SageMaker AI, and the generative model through Amazon Bedrock. Sarvaswa's agent framework coordinates the checks as separate evaluated steps. The same design can be deployed on other clouds where a customer already runs.

Vision is one of the model types we train. See how we fine-tune language models, how the agent framework behind the judging step works, or how we take a build through security review.

Have images or video a person still reviews by hand?

Fifteen minutes is enough to say whether a trained vision model can take the decision on, what it would measure, and what a proof of concept would return on your own examples.