
Quick Summary
- Pfizer’s AI results rest on infrastructure first: it moved 12,000 applications and 8,000 servers to the cloud in 42 weeks, going from 10% to 80% cloud and saving over $47 million a year.
- Centralized data came before models, with the Scientific Data Cloud making instrument data from hundreds of labs searchable and reusable, including failed experiments.
- The clearest gains come from narrow, measured use cases: statistical analysis plans drafted in minutes instead of days, and COVID-19 trial data review-ready in 22 hours instead of 30+ days.
- Generative AI is tightly constrained and human-reviewed, while manufacturing models only advise operators, an approach that cut one Paxlovid step’s cycle time by 67% and added 20,000 doses per batch.
- A federated model gives business teams ownership of AI while a central team runs the platforms, and Pfizer clearly separates its $735 million in measured 2024 impact from targets like $2 billion by 2026.
In 2021, Pfizer moved 12,000 applications and 8,000 servers to the cloud in 42 weeks. It went from 10% cloud to 80% and closed three data centers.
That migration is not the AI story. It is the reason the AI story exists.
By 2024, Pfizer’s AI use cases delivered $735 million in annual impact. Behind that number sit narrower, harder results:
- A statistical analysis plan drafts in one to three minutes instead of one to two days
- COVID-19 trial data was review-ready in 22 hours instead of 30-plus days
- One manufacturing step ran 67% faster and yielded 20,000 more doses per batch
None of these came from buying a better model. Each came from a decision about where AI could act, what it could generate, and who checked its work.
The Problem: Science That Could Not Find Its Own Data
Pharmaceutical development produces information faster than anyone can read it. One drug generates roughly 20,000 documents. Every step is regulated, and every document may be audited.

Before 2019, Pfizer’s scientific data was fragmented in ways that blocked reuse:
- Instrument silos: hundreds of lab instruments produced outputs nobody could search across.
- Manual retrieval: scientists hunted through multiple tools for synthesis routes, formulations, and batch records.
- Late data cleaning: trial data was inspected only after the trial ended, so review waited more than 30 days.
- Slow first drafts: regulatory documents were written from scratch even when most content was reusable.
- Infrastructure drag: only about 10% of core IT ran in the cloud, so compute for big submissions was slow to provision.
Manufacturing had a parallel issue. Continuous production lines generated sensor data that could reveal abnormal equipment behavior. Nobody was looking at it in time to act.
Pfizer had the data. What it lacked was a way to reach it, trust it, and act on it inside the workflows people already used. So it built that access first, then applied AI narrowly and measured it precisely.
What Pfizer Delivered
Pfizer’s annual reviews, AWS case material, and Pfizer-authored papers make the sequence reconstructible. Each step built what the next one needed.
| Period | What Pfizer delivered | Why it mattered |
| 2019 | Scientific Data Cloud with AWS, aggregating hundreds of lab instruments | Made experimental data searchable and reusable |
| 2020 | Smart Data Query cleaning COVID-19 trial data during the trial | Review-ready data in 22 hours instead of 30+ days |
| 2021 | 12,000 applications and 8,000 servers migrated in 42 weeks; PACT formed with AWS | 80% cloud, $47M annual savings, prototype capacity |
| 2021 | Anomaly detection prototype on continuous manufacturing sensors | Early warnings with minimal false positives |
| 2022 | AI-optimized Paxlovid supply chain step; supercomputing for molecule search | 67% shorter cycle time, 20,000 extra doses per batch |
| 2023 | VOX generative AI platform; LLM challenge on safety-table summarization | Plain-language document access; evidence on GenAI limits |
| 2024 | $735M annual AI impact; Medical AI Assistant; demand forecasting for about half of products | AI moved from pilots to enterprise-scale results |
| 2026 | Published GenAI systems for statistical analysis plans and PK reports; federated AI model; 98% AI fluency participation | Controlled generation with evaluation, ownership in the business |
Six capabilities came out of that sequence. Each sits on the one below it.
| Layer | Capability | What it does |
| Infrastructure | 80% cloud, three data centers closed, on-demand compute | Provisions 60,000 CPUs in an hour for large submissions |
| Scientific data | Scientific Data Cloud on an open data-lake architecture | Aggregates and serves instrument data for reuse |
| Knowledge access | VOX with enterprise search and foundation models | Lets scientists query 20,000-document libraries in plain language |
| Clinical data | Smart Data Query | Finds discrepancies while the trial runs |
| Document generation | Statistical analysis plan and PK report drafting systems | Produces controlled first drafts from protocols and study outputs |
| Manufacturing | Anomaly detection, Golden Batch, Digital Operations Center | Warns operators and recommends actions on live production data |
How Pfizer Built It
Six stages. In each one Pfizer had a faster or cheaper option. It chose the one that would hold up under audit.

Stage 1: Move the infrastructure first, then centralize the science
Pfizer’s cloud program started in 2021, mid-pandemic. Most companies would have phased it. Pfizer did it in 42 weeks.
- 12,000 applications and databases migrated
- 8,000 servers migrated
- Cloud footprint from 10% to 80%
- Three data centers closed
- More than $47 million saved every year since
The migration is not the interesting decision. What Pfizer put on top of it is.
Two years earlier, it had built the Scientific Data Cloud with AWS. The problem: hundreds of lab instruments, each writing results in its own format, none searchable together.
The Scientific Data Cloud pulled that data into one open data-lake architecture. Scientists could now find and reuse experiments across the company, including the failed ones.
| Foundation decision | Alternative Pfizer rejected | Consequence |
| Migrate 80% of core IT in under a year | Multi-year phased migration | 60,000 CPUs provisioned in an hour for large submissions |
| Aggregate instrument data centrally | Leave data in instrument-specific systems | Prior experiments reusable, including failed ones |
| Keep the data-lake architecture open | Adopt a closed analytics suite | New models plug into the same data |
Pfizer’s CEO later called 177 years of accumulated data the company’s “alpha.” That only works if the data can be reached.
Stage 2: Clean trial data while the trial is running
Every clinical trial used to end the same way. Data came in, then data scientists spent a month hunting for errors before analysis could start.
For the COVID-19 vaccine trial, Pfizer refused to wait. It deployed Smart Data Query, a machine learning tool built with an external partner that finds discrepancies as data arrives.
The build set a pattern Pfizer would reuse:
- Secure sandbox: develop on anonymized data in an experimentation environment.
- Human in the loop: a reviewer approves or rejects every predicted discrepancy.
- Feedback training: each decision improves the next prediction.
- Operational fit: adapt the tool to how monitoring teams actually work.
The result: data was review-ready 22 hours after the primary efficacy case count. The 44,000-participant trial ran on data cleaned continuously, not at the end.
Stage 3: Prototype with a partner, decide fast, and kill what is not ready
In 2021 Pfizer created the Pfizer-Amazon Collaboration Team, or PACT. It is a delivery model, not a product.
The loop is short:
- Business and technology teams pick a problem by expected value
- AWS engineers build a prototype
- Pfizer decides: minimum viable product, production, or stop
The PACT case study gives the throughput.
| PACT delivery | Figure |
| Projects pursued | 14 |
| Moved into production | 5 |
| Typical prototype duration | No more than 6 weeks |
| Internal estimate for the same work | At least 3 months |
| Estimated scientist time saved from search | Up to 16,000 hours per year |
The best-known PACT product is VOX. One drug can generate 20,000 documents, and scientists used to search them by hand across several tools.
With VOX they ask a question in plain language, typed or spoken. The answer comes from synthesis routes, formulations, analytical tests, and batch records.
Two things about PACT matter more than the tooling:
- Not every prototype shipped. Some were stopped because users were not ready, and Pfizer counted that as a valid outcome.
- The savings figure is an estimate. Pfizer says scientists could save up to 16,000 hours a year, but has not published a measured number.
Stage 4: Constrain generation before scaling it
This is where Pfizer’s discipline shows most. And it started with a failure.
In late 2023 Pfizer invited six external teams to generate safety-table summaries for clinical study reports using LLMs. The test set was 72 real reports across 17 drug assets, judged by blinded expert review.
The results were uneven. The weak spots were the ones a regulator cares about most:
- Factual accuracy
- Concise scientific writing
Fluent output was never the problem. Correct output was.
So Pfizer changed the question. Instead of asking a model to write a document, it asked the model to assemble one.
The statistical analysis plan system published in March 2026 shows how:
- Structured source: protocols parsed into a knowledge graph and vector database.
- Template mapping: protocol content mapped to study-specific templates.
- Four generation modes: copy approved text, summarize, insert variables into standard text, or generate new text only where nothing reusable exists.
- Familiar tools: a web app and a Microsoft Word add-in, so statisticians never leave their normal environment.
| Statistical analysis plan drafting | Figure |
| Plans evaluated | 71 |
| Trial types | Interventional, non-interventional, clinical pharmacology, oncology |
| First draft before | 1 to 2 days |
| First draft after | 1.0 to 3.4 minutes |
| Expert reviewer rating | 3.6 to 4.2 out of 5 |
That rating is not an accuracy score. The system writes the first draft; the statistician still owns the document.
A second system, PK-CSR.Gen, applies the same discipline to pharmacokinetic sections of study reports:
- Chained LLM steps instead of one big prompt
- Fewer than a dozen example reports, no large-scale fine-tuning
- Optional checkpoints where a human steps in
- About 90% of expert-written reporting quality in blinded review
Ninety percent of expert quality is a feasibility result, not a rollout figure. Pfizer says so itself.
Stage 5: Put models next to operators, not in charge of the line
Pfizer’s manufacturing AI follows one rule. The model detects and recommends. The operator decides.
The first documented system came from the 2021 AWS collaboration. Pfizer makes solid oral-dose medicines on continuous lines covered in sensors, and the prototype watched that data for abnormal equipment behavior.
The loop is deliberately incomplete:
- Equipment emits data: sensors on the line stream readings.
- Model detects: anomaly detection flags a deviation.
- Operator is warned: the system shows the signal and its likely source.
- Human intervenes: maintenance or process change is a person’s call.
The measured result comes from Paxlovid. Pfizer teams analyzed supply chain data to find and fix production issues.
In one critical step, cycle time fell 67%. That freed 20,000 extra doses per batch. One step, not the whole line.
Later initiatives keep the same shape:
- Golden Batch: learns desirable process parameters, flags deviations, recommends actions. Yield and cycle-time gains are stated targets.
- Digital Operations Center: end-to-end manufacturing view with real-time collaboration. Pfizer reports 20% higher throughput alongside it, baseline not published.
- Demand forecasting: by 2024, AI predicted demand for about half of Pfizer’s products, with a digital assistant flagging potential shortages early.
Stage 6: Federate ownership, centralize the platform
The last decision is organizational. Pfizer’s CEO laid it out in How Pfizer Thinks About AI.
He chose a federated model over a central AI team. Whoever runs an experiment or a production line knows that work better than any central group.
The split is clean:
- Center owns: platforms, data, computing power, cloud
- Business leaders own: changing how their teams work
- Every colleague owns: AI fluency, through a role-specific certification program launched in 2025
By mid-2026, 98% of eligible colleagues had taken the first course. That is a participation figure, not proof of adoption. But it explains how a Word add-in ends up in the hands of the statisticians who own the documents.
| Operating decision | What it produced |
| Central platforms, local ownership | Business teams change workflows; IT does not have to |
| Role-specific AI fluency for every eligible colleague | Users who can evaluate model output, not just accept it |
| Published responsible-AI principles | Human oversight, privacy, and accountability as design constraints |
| Enterprise impact goal of $2 billion by 2026 | A target that every AI use case is tracked against |
Business Outcomes in Numbers
Pfizer publishes its AI results with unusual precision about scope. A 67% gain on one manufacturing step is a real number. A 67% gain in manufacturing would not be.

The tables below keep each figure attached to the system and period it describes.
Research and clinical development
The clearest measured gains sit where AI handles structured, repetitive work inside a defined document or dataset.
| Metric | Result |
| Statistical analysis plan first draft | 1.0 to 3.4 minutes, from 1 to 2 days, across 71 plans |
| Expert rating of generated plans | 3.6 to 4.2 out of 5 |
| PK study report sections | About 90% of expert-written quality in blinded review |
| COVID-19 trial data ready for review | 22 hours after primary efficacy count, from 30+ days |
| Trial data quality checks and analysis | 50% faster, with AI in more than half of trials by 2022 |
| Manuscript drafts with Medical AI Assistant | 40% less first-draft time, 15% less total submission time |
| Certain scientific calculations | 80 to 90% less computational time |
The drafting numbers describe first drafts. Approval and regulatory submission still take the time they take. What changed is where the statistician’s hours go.
Manufacturing and supply
Manufacturing results are fewer and more tightly scoped. That is appropriate for regulated production.
| Metric | Result |
| Paxlovid critical step cycle time | Down 67% |
| Doses per batch from that step | 20,000 additional |
| Throughput with Digital Operations Center | Up 20%, baseline not disclosed |
| Products with AI demand forecasting | About half of the portfolio, 2024 |
| Anomaly detection on continuous lines | Early warnings with minimal false positives |
Platform and enterprise
Infrastructure results are large. They should not be counted as AI savings.
| Metric | Result |
| Applications and databases migrated | 12,000 in 42 weeks |
| Servers migrated | 8,000 |
| Cloud footprint | From 10% to 80% |
| Annual infrastructure savings | More than $47 million |
| Data centers closed | 3 |
| Compute provisioning | 60,000 CPUs in one hour |
| AI use case impact, 2024 | $735 million annual impact |
| PACT prototype cycle | 6 weeks versus at least 3 months internally |
What is a target, not a result
Several widely quoted Pfizer AI numbers are projections. They belong in their own table.
| Figure | Status |
| 16,000 scientist hours saved per year | Estimated potential, not measured |
| 55% lower infrastructure costs | Estimate in the same case study |
| Golden Batch: 10% higher yield, 25% better cycle time | Stated targets |
| $750 million to $1 billion in GenAI value | 2023 estimate of near-term potential |
| $2 billion in AI impact by 2026 | Goal stated in the 2024 review |
| OncoScout: 25 to 30% experimental success improvement | Target, not reported result |
Pfizer draws these distinctions in its own disclosures. An enterprise reading the case should do the same.
Conclusion
Pfizer’s AI story is usually told as a list of partnerships and one headline number. Read as a sequence of decisions, it is more useful. Pfizer chose to:
- Migrate 80% of core IT in 42 weeks rather than phase it over years
- Aggregate instrument data into a scientific data cloud before building models on it
- Clean trial data during the trial instead of after it
- Prototype in six-week cycles with a partner, and let some prototypes die
- Run a public challenge that exposed GenAI’s accuracy problems before deploying it
- Constrain generation to four modes, delivered inside Microsoft Word
- Keep manufacturing models advisory, with operators in control
- Federate ownership to the people who run experiments and production lines
Each decision produced the data, the evidence, or the trust the next one needed. That is why Pfizer’s published results are narrow and defensible. It is also why the company labels its own projections as projections.
For leaders in regulated industries, the transferable lesson is the order of operations and the discipline of scope. The same pattern underpins legacy modernization and production AI programs that survive audit: infrastructure, then data, then narrowly scoped models with evaluation, then ownership in the business.
The question worth asking is which of those layers is still missing. And whether the next AI investment will land on something that can carry it.
Disclaimer: This is an independent case study review, not a Zoolatech project. It is based on publicly available sources, including Pfizer’s reports, AWS case studies, and published research.
Questions You May Have
How does Pfizer use AI in drug development?
Pfizer applies AI to narrow, well-defined tasks: cleaning clinical trial data while the trial runs, drafting statistical analysis plans and pharmacokinetic report sections, and letting scientists search large document libraries in plain language through its VOX platform. In each case, experts review and own the final output.
Why was the cloud migration so important to Pfizer's AI results?
Moving 80% of core IT to the cloud in 42 weeks gave Pfizer on-demand compute and a foundation for its Scientific Data Cloud. Without centralized, searchable data and scalable infrastructure, later AI tools would have had nothing reliable to run on.
How does Pfizer keep generative AI accurate in regulated documents?
After an LLM challenge showed weaknesses in factual accuracy, Pfizer constrained generation. Its drafting system copies approved text, summarizes, or inserts variables into standard text, and only generates new text where nothing reusable exists. Statisticians review every draft inside familiar tools like Microsoft Word.
What role does AI play in Pfizer's manufacturing?
AI is advisory, not autonomous. Models detect anomalies in sensor data and recommend actions, while operators make the final call. Using this approach, one critical Paxlovid production step ran 67% faster and yielded 20,000 extra doses per batch.
How much value has AI delivered for Pfizer?
Pfizer reported $735 million in annual impact from AI use cases in 2024. Several other widely quoted figures, such as the $2 billion goal for 2026 and 16,000 saved scientist hours, are targets or estimates rather than measured results.










