Training data
The examples a machine-learning model learns from; their quality and coverage shape everything it can do.
At a glance
Key signals
0% cross fields · reaches 5 more
- Machine Learning
- Computer Science
- Statistics
- Explanation
- Examples
- Misconception
- Sourced relations 2
- Attribution
Dependencies
What this concept builds on and what it makes possible — derived from the atlas’s dependency, causal and structural relations, not from every related edge.
Enables · leads to
Gradient descentdepends onTraining dataEstablished
gradient descent depends on Training data.
Mechanism: Each step of gradient descent measures the model's error on training data and adjusts to do better next time.
- Wikipedia (English & German editions) verifiedmoderate evidence
Training dataenablesMachine learningEstablished
Training data is what a model learns from.
Mechanism: A model has nothing to learn until it is given examples; the training data is the raw material of learning.
Model biasis derived fromTraining dataStrongly supported
A model's bias is largely inherited from skewed training data.
Mechanism: Skew in the examples becomes skew in the model: it learns the imbalance as if it were the truth.
Overfittingdepends onTraining dataEstablished
Overfitting is clinging too tightly to the training data.
Mechanism: Overfitting happens when a model memorises the particular training examples instead of the pattern behind them.
Supervised learningdepends onTraining dataEstablished
supervised learning depends on Training data.
Mechanism: Supervised learning learns from labelled training data — examples paired with the right answer — then generalises to new cases.
- Wikipedia (English & German editions) verifiedmoderate evidence
Structural role & consequence
Interpreted from the current atlas graph — what the connections mean, not just how many there are.
Directly enables 1 concept; following enables/causes relations, 3 concepts are downstream across 7 disciplines.
structural · Follows only enables/causes dependency edges — not general relatedness.
Currently dark in the atlas: no key date stored · 5 of 5 of its relations lack claim-level evidence.
atlas representation · Describes the current Thinking OS representation, not the state of the world.
Structural neighbourhood: 5 → 35 → 147 concepts reachable within 3 hops.
structural · Structural reach — being reachable is not the same as being understood.
All 5 of its relationships stay within its own discipline — a field-specific concept in the current atlas.
structural · Structural graph analysis — not a claim of importance, causation or history.
cross-field
5 within-field, 0 cross-field
Strengths & constraints
Constraints
- Evidence coverage currently thin in the atlas — few of its relationships carry claim-level evidence. atlas representation
- No dated history stored — the atlas records no key date for this concept. atlas representation
Conditions
- Read structurally — most of its relationships carry no external evidence yet, so claims here are graph-derived. structural
Dependency radial
What this concept builds on (left) and what it makes possible (right) — derived from dependency and causal relations.
What builds on this
3 concepts build on this directly, 4 in total, across 5 disciplines.
Structural downstream reach along dependency edges — not a claim of historical necessity.
Seen through each discipline
How this concept sits in each of its fields — derived from its real connections in the graph, not asserted.
Through this lens it connects to Machine learning, Supervised learning, Overfitting and Gradient descent.
Through this lens it connects to Machine learning, Supervised learning and Gradient descent.
Through this lens it connects to Machine learning, Overfitting and Model bias.
Check yourself
A quick check against a common misconception. Nothing is scored — picking the tempting-but-wrong answer just flags an idea worth revisiting.
Which statement is correct?
More data always makes a model better.
Only if the data is relevant and representative. Noisy or skewed data just teaches the wrong lesson faster.
Look for: Learner proposes 'add more data' as the fix for every model problem, without asking whether the data is relevant or representative.
Where you'll meet it
Journeys that walk you through this idea. You may recognise it from more than one.
Related ideas to explore
Concepts that look related but are not yet connected here — candidates for a connection to reason about, not established links.
This idea also appears in…
The same structure shows up in other disciplines. These are real recurrences drawn from the graph — a starting point for asking “what carries over, and what changes?”
Information35 disciplines · 29 concepts
Concepts
shares a mental model
shares a mental model
shares a mental model
shares a mental model
shares a mental model
shares a mental model
A learning machine starts with no ideas of its own. It studies many examples — the training data — and adjusts itself to fit them. If those examples are narrow or lopsided, so is everything it learns.
Mental models at work here
- Connects 5 other ideas across 3 disciplines.
- A cross-disciplinary bridge — its connections reach into 5 other fields.
- Most of its connections are of the “Dependency” kind.
- It exercises 1 reusable thinking pattern.
- It features in 1 learning journey.
Derived from the graph’s real structure — observations, not a score.