• 4,000 firms
  • Independent
  • Trusted
Save up to 70% on staff

Home » Glossary » Foundation Model

Foundation Model

Definition

Foundation Model

A foundation model is a large AI system trained on vast unlabeled data via self-supervision, then adapted to many tasks through fine-tuning or prompting. They now underpin most generative AI, from chatbots and code copilots to image and video generators.

The term was coined by Stanford’s Center for Research on Foundation Models (CRFM) in 2021 to describe a new class of AI. Instead of building a fresh model per task, teams train one large model once, then adapt it many times.

This shift changes AI economics. Training costs concentrate at a handful of well-funded labs, while adaptation costs collapse to a fraction of what task-specific models once required.

In 2024, most large organizations reported using generative AI in at least one business function.

You use one every day — GPT-4, Claude, Gemini, and Llama are all foundation models. Businesses tap them through APIs to power chatbots, automate document review, and speed up software development.

Key takeaways

  • Foundation models are trained once on huge, unlabeled datasets, then reused for many tasks through fine-tuning or prompting.
  • Training a frontier model costs tens to hundreds of millions of dollars; adaptation costs a tiny fraction of that.
  • The category covers text, code, image, audio, and multimodal systems — GPT-4, Claude, Gemini, Llama, DALL-E, and Stable Diffusion are examples.
  • Enterprises typically consume foundation models through vendor APIs, private deployments, or open-weights releases.
  • Governance frameworks like the NIST AI Risk Management Framework now target them explicitly as general-purpose systems.

How it works

Foundation models are built in two stages. First, engineers pre-train a large neural network — usually a transformer — on trillions of tokens from books, code, images, or the open web. Then downstream teams adapt that base model for narrower use cases.

StageWhat happensCost / Time
Pre-trainingSelf-supervised learning on trillions of tokensWeeks to months; $10M–$100M+
Fine-tuningAdapting weights on task-specific labeled dataHours to days; $1K–$100K
Prompting or RAGNo weight changes; behavior shaped by inputMilliseconds; near-zero marginal cost

The pre-training step is where the heavy lifting happens. A transformer network learns statistical patterns across the corpus, developing what researchers call emergent capabilities.

Nobody explicitly teaches the model to translate French or summarize contracts; those skills arise from scale.

Most enterprises skip pre-training entirely. They call vendor APIs, fine-tune open-weight checkpoints on private data, or ground responses in their own documents through retrieval-augmented generation. Latency, cost, and data privacy usually drive the choice.

Cost matters here. A frontier pre-training run rents tens of thousands of GPUs for weeks, and a $100 million bill is not unusual. Fine-tuning typically costs a few thousand dollars.

AI research engineer inside a hyperscale data center holding a laptop showing a training-run cost dashboard.
How much does a frontier pre-training run cost?

Examples

Foundation models now dominate frontier AI, and examples span major consumer and enterprise products. OpenAI’s GPT, Anthropic’s Claude, Google’s Gemini, and Meta’s Llama each anchor deployments across search, coding, customer support, and content generation.

Product manager at a modern workstation reviewing four LLM chat interfaces powering search, coding, support and content tasks.
Which foundation models power today’s products?

PT-4 (OpenAI, 2023) shipped as a general-purpose language model; Claude 3 (Anthropic, 2024) added long-context reasoning; Gemini 1.5 (Google DeepMind, 2024) pushed multimodal input; and Llama 3 (Meta, 2024) released open weights.

GPT-4 (OpenAI, 2023). The model that popularized foundation-model APIs. OpenAI released GPT-4 in March 2023, and businesses now use it inside customer-service bots, copywriting tools, and coding assistants like GitHub Copilot.

Claude 3 (Anthropic, 2024). Anthropic launched the Claude 3 family in March 2024, positioning it around safety research and longer-context reasoning. Legal and financial firms use Claude for document review and contract summarization.

Llama 3 (Meta, 2024). Meta released Llama 3 weights publicly in April 2024, enabling businesses to run foundation models on private hardware. It became the default choice for privacy-sensitive workloads in banking and healthcare.

Gemini 1.5 (Google DeepMind, 2024). Google DeepMind released Gemini 1.5 in February 2024, native multimodal with a million-token context window. It handles video, images, and code alongside text, well-suited to research and knowledge-work automation.

Stable Diffusion (Stability AI, 2022). The image model that made generative AI mainstream. Its open release in August 2022 sparked the surge in AI-generated marketing assets, product renders, and design workflows now handled by BPO creative teams.

Related terms

Foundation models sit at the center of a wider AI vocabulary. The terms below cover the fields that produce them, the techniques they build on, and the systems that adapt them for real business tasks.

  • Artificial Intelligence: broad field of building machines that mimic human reasoning, of which foundation models are the current state of the art.
  • Machine Learning: the discipline of teaching computers from data, and the technical parent of every foundation model in production today.
  • Generative AI: applications that create new text, images, audio, or code, most of which are built on top of a foundation model.
  • Natural Language Processing: the sub-field concerned with human language, now dominated by transformer-based foundation models trained on trillions of tokens of web text.
  • Automation: the broader practice of removing manual steps from a process, now often powered by a foundation model doing the reasoning underneath.

FAQ

What is a foundation model in AI?

A foundation model is a large-scale AI system pre-trained on broad, unlabeled data via self-supervision. It can be adapted to many downstream tasks, from writing code to generating images, without retraining from scratch. Stanford researchers coined the term in 2021.

How does a foundation model differ from a traditional machine learning model?

Traditional models are built for one job and trained on curated, task-specific data. Foundation models are trained once on huge, general data and reused across many jobs through fine-tuning or prompting. The result: one model handles what used to require dozens.

Are foundation models the same as large language models?

Large language models (LLMs) are the text-focused subset. Image, audio, and multimodal foundation models exist too. DALL-E, Stable Diffusion, and Gemini 1.5 all sit inside the wider category.

How much does it cost to train a foundation model?

Frontier training runs cost tens to hundreds of millions of dollars in compute and data. That’s why only a handful of labs train them from scratch; most enterprises consume foundation models through APIs or open-weights releases.

Which industries use foundation models the most?

Software, finance, healthcare, media, and legal services show the earliest adoption. Companies use them to summarize documents, draft first-pass content, answer customer queries, and automate research. Regulated industries move slower due to data-privacy constraints.

Do foundation models replace BPO work?

No, foundation models augment BPO teams by automating repetitive first-pass work while humans handle exceptions and quality control.

Explore how outsourcing partners can help you deploy foundation-model applications at Outsource Accelerator.

Companies you might be interested in

Get Inside Outsourcing

An insider's view on why remote and offshore staffing is radically changing the future of work.

Order now

Start your
journey today

  • Independent
  • Secure
  • Transparent

About OA

Outsource Accelerator is the trusted source of independent information, advisory and expert implementation of Business Process Outsourcing (BPO).

The #1 outsourcing authority

Outsource Accelerator offers the world’s leading aggregator marketplace for outsourcing. It specifically provides the conduit between world-leading outsourcing suppliers and the businesses – clients – across the globe.

The Outsource Accelerator website has over 5,000 articles, 450+ podcast episodes, and a comprehensive directory with 4,700+ BPO companies… all designed to make it easier for clients to learn about – and engage with – outsourcing.

About Derek Gallimore

Derek Gallimore has been in business for 20 years, outsourcing for over eight years, and has been living in Manila (the heart of global outsourcing) since 2014. Derek is the founder and CEO of Outsource Accelerator, and is regarded as a leading expert on all things outsourcing.

“Excellent service for outsourcing advice and expertise for my business.”

Learn more
Banner Image
Get 3 Free Quotes Verified Outsourcing Suppliers
4,000 firms.Just 2 minutes to complete.
SAVE UP TO
70% ON STAFF COSTS
Learn more

Connect with over 4,000 outsourcing services providers.

Banner Image

Transform your business with skilled offshore talent.

  • 4,000 firms
  • Simple
  • Transparent
Banner Image