For the complete documentation index, see llms.txt. This page is also available as Markdown.

How our multimodal AI works

See how Dstillery connects behavioral, language, and partner signals to understand real audience intent.

People do not move through the internet in channels. They search, watch, read, compare, and return. Dstillery's multimodal AI turns those connected moments into one understanding of intent.

One model. More complete intent. DS-1 reasons from the signals behind connected actions.

See the journey, not the channel

Picture researching a new coffee maker. You Google "best espresso machines," browse a retailer's site, watch a review on YouTube, then see a CTV ad before buying.

Each moment adds context. Together, they reveal intent. Most targeting evaluates those moments separately. Dstillery's model learns how they relate over time.

What multimodal means

Multimodal doesn't just mean more data sources. It means learning across different forms of information.

Words carry meaning. Behaviors form sequences. The model learns from both together, along with everything in between.

Four signals. One intelligence layer.

1

Website visitation data

Billions of daily behavioral events across more than 400 million devices reveal patterns at scale.

2

Opted-in panel data

An opted-in panel of about 2 million users shows the behavior sequences that lead to real decisions, without relying on cookies or IDs.

3

LLM-derived insights

Large language models connect concepts across signals. They can link CTV viewership to web browsing, or turn a written audience description into targetable patterns.

4

Partner data

Specialized signals add depth for CPG, healthcare, B2B, and retail use cases.

The first two signal types here are the same two datasets covered in Dstillery data & DS-1: how they fit together. For the full list of data partnerships behind signal four, see Our data sources.

The shared language: embeddings

Embeddings let different signals work together. They map an advertising opportunity into a shared space of meaning, the same way a language model learns embeddings that capture what words mean.

That shared space connects a search term to a CTV pattern. It can turn a plain-language brief into behavioral signal. The model isn't comparing disconnected systems, it's working in one shared language for intent.

Start anywhere. Activate everywhere.

Multimodal learning lets the model begin with almost any useful input:

  • First-party data or a CRM list

  • A search keyword or URL

  • A plain-language audience description

It activates that same understanding through user segments, contextual targeting, private marketplace deals, and custom bidding.

Built for measurable outcomes

This isn't a marginal improvement. It changes the quality of the signal behind each decision.

An auto insurance brand testing direct response found contextual targeting beat 12 other tactics, including ID-based lookalikes.

The foundation for agentic advertising

Agentic advertising needs more than contextual inference. It needs a complete view of intent.

An AI agent reasoning from contextual signals alone would be missing behavioral signals, which are a much better predictor of intent. Multimodal understanding is what lets an agent like DS-1 reason across every available signal type at once.

Last updated

Was this helpful?