Why you should always build a baseline model first

Why you should always build a baseline model first

Before training a real model, measure the simplest possible one. On a price-like series, “tomorrow = today” beat a random forest by more than 3×. Tested with scikit-learn, with the reason the forest failed.

TL;DR — A baseline is the simplest prediction you could make without machine learning, such as “tomorrow equals today” or “always predict the average”. Measure it first. If your model can’t beat it clearly, the model isn’t learning anything useful. In our test, the naive baseline scored an error of 0.79; a random forest scored 2.55.

Tested on 2026-09-28 with scikit-learn 1.9.1 on a synthetic random walk (1,200 days).

Key points

  • A baseline answers “how good is good?” Without it, an error of 2.55 means nothing.
  • Pick a baseline that fits the problem: the last value for time series, the most common class for classification, the mean for regression.
  • A bad baseline flatters your model. “Training mean” had an error of 34.01 here, which makes almost anything look great.
  • Tree models can’t predict outside the range they saw in training. That’s why the forest lost.
  • A model that only matches the baseline has usually just rediscovered it.

Can a simple rule beat a machine learning model?

Here is a price-like series where each day is yesterday plus a random change (a random walk). We predict the next day from the last five days, training on the first 1,000 days and testing on the rest:

import numpy as np
from sklearn.ensemble import RandomForestRegressor
from sklearn.linear_model import LinearRegression
from sklearn.metrics import mean_absolute_error

# A price-like series: each day = yesterday + random change (a random walk)
rng = np.random.default_rng(0)
price = 100 + np.cumsum(rng.normal(0, 1, 1_200))

# Features: the last 5 days. Target: the next day.
lags = 5
X = np.column_stack([price[i:len(price) - lags + i] for i in range(lags)])
y = price[lags:]
split = 1_000                                   # train on the past, test on the future
X_tr, X_te, y_tr, y_te = X[:split], X[split:], y[:split], y[split:]

results = {
    "baseline: tomorrow = today": mean_absolute_error(y_te, X_te[:, -1]),
    "baseline: training mean":    mean_absolute_error(y_te, np.full_like(y_te, y_tr.mean())),
    "linear regression":          mean_absolute_error(y_te, LinearRegression().fit(X_tr, y_tr).predict(X_te)),
    "random forest":              mean_absolute_error(y_te, RandomForestRegressor(random_state=0).fit(X_tr, y_tr).predict(X_te)),
}
print(f"train range: {y_tr.min():.1f} to {y_tr.max():.1f} | test range: {y_te.min():.1f} to {y_te.max():.1f}")
for name, mae in results.items():
    print(f"{name:<28} MAE {mae:6.2f}")
train range: 50.1 to 111.5 | test range: 39.8 to 65.4
baseline: tomorrow = today   MAE   0.79
baseline: training mean      MAE  34.01
linear regression            MAE   0.80
random forest                MAE   2.55

MAE (mean absolute error) is the average size of the miss, in price units. Lower is better.

ModelMAEVerdict
Baseline: tomorrow = today0.79Best
Linear regression0.80Ties the baseline
Random forest2.55More than 3× worse than the baseline
Baseline: training mean34.01Wrong baseline for time series

Why did the random forest lose?

A random forest predicts by averaging training examples, so it can never predict a value outside the range it saw. The test period went down to 39.8, but the lowest training value was 50.1:

# continues from the code above
lin = LinearRegression().fit(X_tr, y_tr)
print("linear weights on the last 5 days:", lin.coef_.round(2), "| intercept:", round(lin.intercept_, 2))
rf_pred = RandomForestRegressor(random_state=0).fit(X_tr, y_tr).predict(X_te)
print(f"lowest random forest prediction: {rf_pred.min():.1f} (lowest training target: {y_tr.min():.1f}, lowest test value: {y_te.min():.1f})")
linear weights on the last 5 days: [-0.01  0.    0.02 -0.01  1.01] | intercept: -0.11
lowest random forest prediction: 51.2 (lowest training target: 50.1, lowest test value: 39.8)

Line chart of a 1,200-day random walk. In the test period the actual price drops to about 40; the naive baseline follows it closely, while the random forest stays flat near 52 because it cannot predict below the lowest training value of 50.1

  • The forest’s lowest prediction was 51.2, so it stayed flat while the real price fell to 39.8.
  • Linear regression put a weight of 1.01 on today’s value and about 0 on the other days. It learned “tomorrow = today”, the baseline itself.

For a true random walk, “tomorrow = today” is already the best possible forecast, so no model can beat it by much. The baseline tells you that immediately; the forest’s error alone would not.

Which baseline should you use?

ProblemBaselineIn scikit-learn
ClassificationAlways predict the most common classDummyClassifier(strategy="most_frequent")
Imbalanced classificationPredict classes in their training proportionsDummyClassifier(strategy="stratified")
RegressionAlways predict the training mean or medianDummyRegressor(strategy="mean")
Time seriesLast value, or the value one season agoX_te[:, -1], as above

Report the baseline next to every model score. “MAE 2.55” is a number; “MAE 2.55 against a naive 0.79” is a decision.

Sources

Related: Why is 99% accuracy useless? Precision, recall and F1 explained