backed by a16z speedrun

Panorama turns business challenges into AI-powered solutions.

Whether you're shipping your first AI feature or scaling what's already live, we help take it further.

0xLLM cost reduction
0xFaster time to fast result
0wkFrom engagement to shipped system
0M+Users served by systems this team has shipped

How we work

Strategy

We learn how your business and systems work, and find the changes that create the most value. Sometimes that's a model, sometimes it's the data around it, sometimes it's a process nobody's automated yet. You get a clear picture of what's worth building before any code written or system built.

Card Image

Productionize

We take it past the prototype and into your stack - production-grade, with the evaluations, monitoring, and runbooks that keep it improving over time. Built by engineers with a track record shipping to hundreds of millions of users - who know rigor is what makes systems last.

Card Image

Partnership

We work alongside your team, so your engineers understand what we built and can keep extending it long after we're gone. And when things change - as AI systems always do - we're still reachable.

Card Image

Systems we've built

01
Retrieval
Your RAG pipeline works in the demo and misses in production. We rebuild retrieval as a measured system—from chunking and embeddings to reranking, hybrid search, and evaluation.
02
Context engineering
Agents that forget, prompts that spool to 40K tokens, and context windows spent on the wrong things. We design what the model sees at every step.
03
LLM cost optimization
Your inference bill scales faster than your usage. We profile where tokens go, then reduce cost through routing, caching, batching, and smaller models.
04
Data strategy
You have the data, but it is not in a shape a model can use. We build pipelines, training sets, evaluation sets, labeling systems, and governance.
05
Post-training
When prompting stops being enough, we handle data curation, training, evaluation, deployment, and optimization.

Systems we've built

01
Retrieval
Your RAG pipeline works in the demo and misses in production. We rebuild retrieval as a measured system—from chunking and embeddings to reranking, hybrid search, and evaluation.
02
Context engineering
Agents that forget, prompts that spool to 40K tokens, and context windows spent on the wrong things. We design what the model sees at every step.
03
LLM cost optimization
Your inference bill scales faster than your usage. We profile where tokens go, then reduce cost through routing, caching, batching, and smaller models.
04
Data strategy
You have the data, but it is not in a shape a model can use. We build pipelines, training sets, evaluation sets, labeling systems, and governance.
05
Post-training
When prompting stops being enough, we handle data curation, training, evaluation, deployment, and optimization.
01
Retrieval
Your RAG pipeline works in the demo and misses in production. We rebuild retrieval as a measured system—from chunking and embeddings to reranking, hybrid search, and evaluation.
02
Context engineering
Agents that forget, prompts that spool to 40K tokens, and context windows spent on the wrong things. We design what the model sees at every step.
03
LLM cost optimization
Your inference bill scales faster than your usage. We profile where tokens go, then reduce cost through routing, caching, batching, and smaller models.
04
Data strategy
You have the data, but it is not in a shape a model can use. We build pipelines, training sets, evaluation sets, labeling systems, and governance.
05
Post-training
When prompting stops being enough, we handle data curation, training, evaluation, deployment, and optimization.

What we've shipped.

01

We cut the inference cost by 16x

A frontier model was writing a product’s core output at 11 cents a pop. A small open model, trained on a rented GPU for $8, does the same job for less than a cent.

We audited every call the product made to a model, found one repetitive task eating the whole bill, and distilled it into a focused open model.

16×
Cheaper per call
$8
Total training cost
$99 → $6
Monthly at their volume
02

The prompt was too big to send

A product needed to read more conversation than any frontier model would accept in one call. We restructured what the model sees, and it went from not working at all to running behind a cache.

The real constraint was signal density, not context size. We built a four-stage pipeline — normalize, label-and-drop, rank by evidence density, then generate — running ahead of time into an edge cache. Output got sharper as the prompt got smaller.

Didn't fit → 50k
Prompt size, tokens
5 min → cache read
What the user waits for
03

The best people spent the whole day on outreach

A client asked where AI would be worth using in their business. The audit pointed at outreach.

We built an agent that lives in Slack, drafts messages in each person’s voice, and hands every final decision back to a human.

15 → 1 min
Per prospect
90 → 150
Prospects per day
~20 hrs/day
Skilled time handed back

What we've shipped

01

We cut the inference cost by 16×

A frontier model was writing a product’s core output at 11 cents a pop. A small open model, trained on a rented GPU for $8, does the same job for less than a cent.

We audited every call the product made to a model, found one repetitive task eating the whole bill, and distilled it into a focused open model.

16×
Cheaper per call
$8
Total training cost
$99 → $6
Monthly at their volume
02

The prompt was too big to send

A product needed to read more conversation than any frontier model would accept in one call. We restructured what the model sees, and it went from not working at all to running behind a cache.

The real constraint was signal density, not context size. We built a four-stage pipeline — normalize, label-and-drop, rank by evidence density, then generate — running ahead of time into an edge cache. Output got sharper as the prompt got smaller.

Didn't fit → 50k
Prompt size, tokens
5 min → cache read
What the user waits for
03

The best people spent the whole day on outreach

A client asked where AI would be worth using in their business. The audit pointed at outreach.

We built an agent that lives in Slack, drafts messages in each person’s voice, and hands every final decision back to a human.

15 → 1 min
Per prospect
90 → 150
Prospects per day
~20 hrs/day
Skilled time handed back

What we've shipped

01

We cut the inference cost by 16×

A frontier model was writing a product’s core output at 11 cents a pop. A small open model, trained on a rented GPU for $8, does the same job for less than a cent.

We audited every call the product made to a model, found one repetitive task eating the whole bill, and distilled it into a focused open model.

16×
Cheaper per call
$8
Total training cost
$99 → $6
Monthly at their volume
02

The prompt was too big to send

A product needed to read more conversation than any frontier model would accept in one call. We restructured what the model sees, and it went from not working at all to running behind a cache.

The real constraint was signal density, not context size. We built a four-stage pipeline — normalize, label-and-drop, rank by evidence density, then generate — running ahead of time into an edge cache. Output got sharper as the prompt got smaller.

Didn't fit → 50k
Prompt size, tokens
5 min → cache read
What the user waits for
03

The best people spent the whole day on outreach

A client asked where AI would be worth using in their business. The audit pointed at outreach.

We built an agent that lives in Slack, drafts messages in each person’s voice, and hands every final decision back to a human.

15 → 1 min
Per prospect
90 → 150
Prospects per day
~20 hrs/day
Skilled time handed back

Our team

Our team

Our team

Panorama is a small team of engineers from Google, Twitter, and UC Berkeley, who have shipped systems for hundreds of millions of users, and led high-profile open-source and special projects at Lyft and Cash App. We combine deep customer empathy with a rare ability to turn state-of-the-art technology into working products. We take on real-world problems and turn them into solid systems.

Team member 1
Team member 2
Team member 3

Panorama is a small team of engineers from Google, Twitter, and UC Berkeley, who have shipped systems for hundreds of millions of users, and led high-profile open-source and special projects at Lyft and Cash App. We combine deep customer empathy with a rare ability to turn state-of-the-art technology into working products. We take on real-world problems and turn them into solid systems.

Team member 1
Team member 2
Team member 3

Blog

Blog

Blog

Latest Posts

Backed by top builders from