Pfizer case study

Quick Summary

Key takeaways from the article
  • Pfizer’s AI results rest on infrastructure first: it moved 12,000 applications and 8,000 servers to the cloud in 42 weeks, going from 10% to 80% cloud and saving over $47 million a year.
  • Centralized data came before models, with the Scientific Data Cloud making instrument data from hundreds of labs searchable and reusable, including failed experiments.
  • The clearest gains come from narrow, measured use cases: statistical analysis plans drafted in minutes instead of days, and COVID-19 trial data review-ready in 22 hours instead of 30+ days.
  • Generative AI is tightly constrained and human-reviewed, while manufacturing models only advise operators, an approach that cut one Paxlovid step’s cycle time by 67% and added 20,000 doses per batch.
  • A federated model gives business teams ownership of AI while a central team runs the platforms, and Pfizer clearly separates its $735 million in measured 2024 impact from targets like $2 billion by 2026.

In 2021, Pfizer moved 12,000 applications and 8,000 servers to the cloud in 42 weeks. It went from 10% cloud to 80% and closed three data centers.

That migration is not the AI story. It is the reason the AI story exists.

By 2024, Pfizer’s AI use cases delivered $735 million in annual impact. Behind that number sit narrower, harder results:

  • A statistical analysis plan drafts in one to three minutes instead of one to two days
  • COVID-19 trial data was review-ready in 22 hours instead of 30-plus days
  • One manufacturing step ran 67% faster and yielded 20,000 more doses per batch

None of these came from buying a better model. Each came from a decision about where AI could act, what it could generate, and who checked its work.

The Problem: Science That Could Not Find Its Own Data

Pharmaceutical development produces information faster than anyone can read it. One drug generates roughly 20,000 documents. Every step is regulated, and every document may be audited.

Pfizer case study

Before 2019, Pfizer’s scientific data was fragmented in ways that blocked reuse:

  • Instrument silos: hundreds of lab instruments produced outputs nobody could search across.
  • Manual retrieval: scientists hunted through multiple tools for synthesis routes, formulations, and batch records.
  • Late data cleaning: trial data was inspected only after the trial ended, so review waited more than 30 days.
  • Slow first drafts: regulatory documents were written from scratch even when most content was reusable.
  • Infrastructure drag: only about 10% of core IT ran in the cloud, so compute for big submissions was slow to provision.

Manufacturing had a parallel issue. Continuous production lines generated sensor data that could reveal abnormal equipment behavior. Nobody was looking at it in time to act.

Pfizer had the data. What it lacked was a way to reach it, trust it, and act on it inside the workflows people already used. So it built that access first, then applied AI narrowly and measured it precisely.

What Pfizer Delivered

Pfizer’s annual reviews, AWS case material, and Pfizer-authored papers make the sequence reconstructible. Each step built what the next one needed.

PeriodWhat Pfizer deliveredWhy it mattered
2019Scientific Data Cloud with AWS, aggregating hundreds of lab instrumentsMade experimental data searchable and reusable
2020Smart Data Query cleaning COVID-19 trial data during the trialReview-ready data in 22 hours instead of 30+ days
202112,000 applications and 8,000 servers migrated in 42 weeks; PACT formed with AWS80% cloud, $47M annual savings, prototype capacity
2021Anomaly detection prototype on continuous manufacturing sensorsEarly warnings with minimal false positives
2022AI-optimized Paxlovid supply chain step; supercomputing for molecule search67% shorter cycle time, 20,000 extra doses per batch
2023VOX generative AI platform; LLM challenge on safety-table summarizationPlain-language document access; evidence on GenAI limits
2024$735M annual AI impact; Medical AI Assistant; demand forecasting for about half of productsAI moved from pilots to enterprise-scale results
2026Published GenAI systems for statistical analysis plans and PK reports; federated AI model; 98% AI fluency participationControlled generation with evaluation, ownership in the business

Six capabilities came out of that sequence. Each sits on the one below it.

LayerCapabilityWhat it does
Infrastructure80% cloud, three data centers closed, on-demand computeProvisions 60,000 CPUs in an hour for large submissions
Scientific dataScientific Data Cloud on an open data-lake architectureAggregates and serves instrument data for reuse
Knowledge accessVOX with enterprise search and foundation modelsLets scientists query 20,000-document libraries in plain language
Clinical dataSmart Data QueryFinds discrepancies while the trial runs
Document generationStatistical analysis plan and PK report drafting systemsProduces controlled first drafts from protocols and study outputs
ManufacturingAnomaly detection, Golden Batch, Digital Operations CenterWarns operators and recommends actions on live production data

How Pfizer Built It

Six stages. In each one Pfizer had a faster or cheaper option. It chose the one that would hold up under audit.

Pfizer case study

Stage 1: Move the infrastructure first, then centralize the science

Pfizer’s cloud program started in 2021, mid-pandemic. Most companies would have phased it. Pfizer did it in 42 weeks.

  • 12,000 applications and databases migrated
  • 8,000 servers migrated
  • Cloud footprint from 10% to 80%
  • Three data centers closed
  • More than $47 million saved every year since

The migration is not the interesting decision. What Pfizer put on top of it is.

Two years earlier, it had built the Scientific Data Cloud with AWS. The problem: hundreds of lab instruments, each writing results in its own format, none searchable together.

The Scientific Data Cloud pulled that data into one open data-lake architecture. Scientists could now find and reuse experiments across the company, including the failed ones.

Foundation decisionAlternative Pfizer rejectedConsequence
Migrate 80% of core IT in under a yearMulti-year phased migration60,000 CPUs provisioned in an hour for large submissions
Aggregate instrument data centrallyLeave data in instrument-specific systemsPrior experiments reusable, including failed ones
Keep the data-lake architecture openAdopt a closed analytics suiteNew models plug into the same data

Pfizer’s CEO later called 177 years of accumulated data the company’s “alpha.” That only works if the data can be reached.

Stage 2: Clean trial data while the trial is running

Every clinical trial used to end the same way. Data came in, then data scientists spent a month hunting for errors before analysis could start.

For the COVID-19 vaccine trial, Pfizer refused to wait. It deployed Smart Data Query, a machine learning tool built with an external partner that finds discrepancies as data arrives.

The build set a pattern Pfizer would reuse:

  1. Secure sandbox: develop on anonymized data in an experimentation environment.
  2. Human in the loop: a reviewer approves or rejects every predicted discrepancy.
  3. Feedback training: each decision improves the next prediction.
  4. Operational fit: adapt the tool to how monitoring teams actually work.

The result: data was review-ready 22 hours after the primary efficacy case count. The 44,000-participant trial ran on data cleaned continuously, not at the end.

Stage 3: Prototype with a partner, decide fast, and kill what is not ready

In 2021 Pfizer created the Pfizer-Amazon Collaboration Team, or PACT. It is a delivery model, not a product.

The loop is short:

  • Business and technology teams pick a problem by expected value
  • AWS engineers build a prototype
  • Pfizer decides: minimum viable product, production, or stop

The PACT case study gives the throughput.

PACT deliveryFigure
Projects pursued14
Moved into production5
Typical prototype durationNo more than 6 weeks
Internal estimate for the same workAt least 3 months
Estimated scientist time saved from searchUp to 16,000 hours per year

The best-known PACT product is VOX. One drug can generate 20,000 documents, and scientists used to search them by hand across several tools.

With VOX they ask a question in plain language, typed or spoken. The answer comes from synthesis routes, formulations, analytical tests, and batch records.

Two things about PACT matter more than the tooling:

  • Not every prototype shipped. Some were stopped because users were not ready, and Pfizer counted that as a valid outcome.
  • The savings figure is an estimate. Pfizer says scientists could save up to 16,000 hours a year, but has not published a measured number.

Stage 4: Constrain generation before scaling it

This is where Pfizer’s discipline shows most. And it started with a failure.

In late 2023 Pfizer invited six external teams to generate safety-table summaries for clinical study reports using LLMs. The test set was 72 real reports across 17 drug assets, judged by blinded expert review.

The results were uneven. The weak spots were the ones a regulator cares about most:

  • Factual accuracy
  • Concise scientific writing

Fluent output was never the problem. Correct output was.

So Pfizer changed the question. Instead of asking a model to write a document, it asked the model to assemble one.

The statistical analysis plan system published in March 2026 shows how:

  • Structured source: protocols parsed into a knowledge graph and vector database.
  • Template mapping: protocol content mapped to study-specific templates.
  • Four generation modes: copy approved text, summarize, insert variables into standard text, or generate new text only where nothing reusable exists.
  • Familiar tools: a web app and a Microsoft Word add-in, so statisticians never leave their normal environment.
Statistical analysis plan draftingFigure
Plans evaluated71
Trial typesInterventional, non-interventional, clinical pharmacology, oncology
First draft before1 to 2 days
First draft after1.0 to 3.4 minutes
Expert reviewer rating3.6 to 4.2 out of 5

That rating is not an accuracy score. The system writes the first draft; the statistician still owns the document.

A second system, PK-CSR.Gen, applies the same discipline to pharmacokinetic sections of study reports:

  • Chained LLM steps instead of one big prompt
  • Fewer than a dozen example reports, no large-scale fine-tuning
  • Optional checkpoints where a human steps in
  • About 90% of expert-written reporting quality in blinded review

Ninety percent of expert quality is a feasibility result, not a rollout figure. Pfizer says so itself.

Stage 5: Put models next to operators, not in charge of the line

Pfizer’s manufacturing AI follows one rule. The model detects and recommends. The operator decides.

The first documented system came from the 2021 AWS collaboration. Pfizer makes solid oral-dose medicines on continuous lines covered in sensors, and the prototype watched that data for abnormal equipment behavior.

The loop is deliberately incomplete:

  1. Equipment emits data: sensors on the line stream readings.
  2. Model detects: anomaly detection flags a deviation.
  3. Operator is warned: the system shows the signal and its likely source.
  4. Human intervenes: maintenance or process change is a person’s call.

The measured result comes from Paxlovid. Pfizer teams analyzed supply chain data to find and fix production issues.

In one critical step, cycle time fell 67%. That freed 20,000 extra doses per batch. One step, not the whole line.

Later initiatives keep the same shape:

  • Golden Batch: learns desirable process parameters, flags deviations, recommends actions. Yield and cycle-time gains are stated targets.
  • Digital Operations Center: end-to-end manufacturing view with real-time collaboration. Pfizer reports 20% higher throughput alongside it, baseline not published.
  • Demand forecasting: by 2024, AI predicted demand for about half of Pfizer’s products, with a digital assistant flagging potential shortages early.

Stage 6: Federate ownership, centralize the platform

The last decision is organizational. Pfizer’s CEO laid it out in How Pfizer Thinks About AI.

He chose a federated model over a central AI team. Whoever runs an experiment or a production line knows that work better than any central group.

The split is clean:

  • Center owns: platforms, data, computing power, cloud
  • Business leaders own: changing how their teams work
  • Every colleague owns: AI fluency, through a role-specific certification program launched in 2025

By mid-2026, 98% of eligible colleagues had taken the first course. That is a participation figure, not proof of adoption. But it explains how a Word add-in ends up in the hands of the statisticians who own the documents.

Operating decisionWhat it produced
Central platforms, local ownershipBusiness teams change workflows; IT does not have to
Role-specific AI fluency for every eligible colleagueUsers who can evaluate model output, not just accept it
Published responsible-AI principlesHuman oversight, privacy, and accountability as design constraints
Enterprise impact goal of $2 billion by 2026A target that every AI use case is tracked against

Business Outcomes in Numbers

Pfizer publishes its AI results with unusual precision about scope. A 67% gain on one manufacturing step is a real number. A 67% gain in manufacturing would not be.

Pfizer case study

The tables below keep each figure attached to the system and period it describes.

Research and clinical development

The clearest measured gains sit where AI handles structured, repetitive work inside a defined document or dataset.

MetricResult
Statistical analysis plan first draft1.0 to 3.4 minutes, from 1 to 2 days, across 71 plans
Expert rating of generated plans3.6 to 4.2 out of 5
PK study report sectionsAbout 90% of expert-written quality in blinded review
COVID-19 trial data ready for review22 hours after primary efficacy count, from 30+ days
Trial data quality checks and analysis50% faster, with AI in more than half of trials by 2022
Manuscript drafts with Medical AI Assistant40% less first-draft time, 15% less total submission time
Certain scientific calculations80 to 90% less computational time

The drafting numbers describe first drafts. Approval and regulatory submission still take the time they take. What changed is where the statistician’s hours go.

Manufacturing and supply

Manufacturing results are fewer and more tightly scoped. That is appropriate for regulated production.

MetricResult
Paxlovid critical step cycle timeDown 67%
Doses per batch from that step20,000 additional
Throughput with Digital Operations CenterUp 20%, baseline not disclosed
Products with AI demand forecastingAbout half of the portfolio, 2024
Anomaly detection on continuous linesEarly warnings with minimal false positives

Platform and enterprise

Infrastructure results are large. They should not be counted as AI savings.

MetricResult
Applications and databases migrated12,000 in 42 weeks
Servers migrated8,000
Cloud footprintFrom 10% to 80%
Annual infrastructure savingsMore than $47 million
Data centers closed3
Compute provisioning60,000 CPUs in one hour
AI use case impact, 2024$735 million annual impact
PACT prototype cycle6 weeks versus at least 3 months internally

What is a target, not a result

Several widely quoted Pfizer AI numbers are projections. They belong in their own table.

FigureStatus
16,000 scientist hours saved per yearEstimated potential, not measured
55% lower infrastructure costsEstimate in the same case study
Golden Batch: 10% higher yield, 25% better cycle timeStated targets
$750 million to $1 billion in GenAI value2023 estimate of near-term potential
$2 billion in AI impact by 2026Goal stated in the 2024 review
OncoScout: 25 to 30% experimental success improvementTarget, not reported result

Pfizer draws these distinctions in its own disclosures. An enterprise reading the case should do the same.

Conclusion

Pfizer’s AI story is usually told as a list of partnerships and one headline number. Read as a sequence of decisions, it is more useful. Pfizer chose to:

  • Migrate 80% of core IT in 42 weeks rather than phase it over years
  • Aggregate instrument data into a scientific data cloud before building models on it
  • Clean trial data during the trial instead of after it
  • Prototype in six-week cycles with a partner, and let some prototypes die
  • Run a public challenge that exposed GenAI’s accuracy problems before deploying it
  • Constrain generation to four modes, delivered inside Microsoft Word
  • Keep manufacturing models advisory, with operators in control
  • Federate ownership to the people who run experiments and production lines

Each decision produced the data, the evidence, or the trust the next one needed. That is why Pfizer’s published results are narrow and defensible. It is also why the company labels its own projections as projections.

For leaders in regulated industries, the transferable lesson is the order of operations and the discipline of scope. The same pattern underpins legacy modernization and production AI programs that survive audit: infrastructure, then data, then narrowly scoped models with evaluation, then ownership in the business.

The question worth asking is which of those layers is still missing. And whether the next AI investment will land on something that can carry it.

Disclaimer: This is an independent case study review, not a Zoolatech project. It is based on publicly available sources, including Pfizer’s reports, AWS case studies, and published research.