Skip to content

Feature Engineering Intermediate

📊 Data & AI
⏱️ ~4 days 📚 Prerequisites: Data Cleaning, Statistics

When you'd use this

Turn raw data into features that make models work.

Turn raw data into model-ready features — encoding, scaling, binning, extraction — often the biggest lever on ML model quality.

What you'll learn

  • Why feature engineering matters most
  • Encoding categorical data (tested)
  • Scaling and normalization (tested)
  • Feature creation and selection
  • The scikit-learn workflow

Feature engineering is transforming raw data into inputs (features) that help a model learn. It's often said that feature engineering beats algorithm choice — good features with a simple model usually outperform poor features with a fancy one. The transforms here are run-verified in pure Python.


Encoding categorical data (tested)

Encoding categorical data in Feature Engineering — what it is and when to use it.

Models need numbers, but data has categories ("red", "blue"). One-hot encoding turns each category into its own 0/1 column, avoiding a false ordering. Runnable:

def one_hot(values):
    categories = sorted(set(values))
    return [{c: (1 if v == c else 0) for c in categories} for v in values]

for row in one_hot(["red", "blue", "red"]):
    print(row)

Output:

{'blue': 0, 'red': 1}
{'blue': 1, 'red': 0}
{'blue': 0, 'red': 1}

Each color becomes a pair of 0/1 columns. Why one-hot and not just "red=1, blue=2"? Because numbering categories invents a false order and distance (it would imply blue is "twice" red), which misleads the model. One-hot avoids that. (The tradeoff: many categories → many columns; for high-cardinality data you'd use target/embedding encoding instead.)


Scaling and normalization (tested)

Scaling and normalization in Feature Engineering — what it is and when to use it.

Features on different scales (age 0-100 vs income 0-1,000,000) can bias models that use distances or gradients. Two standard fixes:

Min-max scaling — squash to [0, 1]:

def min_max_scale(values):
    lo, hi = min(values), max(values)
    return [(v - lo) / (hi - lo) for v in values]

print(min_max_scale([10, 20, 30]))

Output:

[0.0, 0.5, 1.0]

Standardization (z-score) — center at 0 with unit standard deviation:

def standardize(values):
    mean = sum(values) / len(values)
    var = sum((v - mean) ** 2 for v in values) / len(values)
    std = var ** 0.5
    return [(v - mean) / std for v in values]

print([round(z, 3) for z in standardize([2, 4, 6])])

Output:

[-1.225, 0.0, 1.225]

Standardized data has mean 0 and std 1. When to use which: min-max keeps values bounded (good for neural nets, image pixels); standardization handles outliers better and suits distance/gradient methods (SVM, linear/logistic regression, k-means). Many models (tree-based ones) don't need scaling at all.


Creating features

Derive new signals (ratios, dates, interactions) that help models learn.

Often the biggest wins come from creating features from domain knowledge:

  • Combinations — price_per_sqft = price / area; ratios and interactions often capture what raw columns don't.
  • Datetime parts — extract day-of-week, month, is_weekend, hour from a timestamp (behavior varies by these).
  • Binning — group a continuous value into ranges (age → "child/adult/senior").
  • Aggregations — per-user averages, counts, recency.
  • Text — word counts, TF-IDF, embeddings (see the AI & LLMs section).

This is where domain understanding turns into predictive power — a model can only learn from what you give it.


Feature selection

Keep the informative features and drop noise to improve and simplify models.

More features isn't always better — irrelevant ones add noise and overfitting risk. Selection keeps the useful ones:

  • Filter — rank features by correlation/statistical test with the target, keep the top.
  • Wrapper — try subsets, measure model performance (e.g. recursive feature elimination).
  • Embedded — the model selects during training (L1/Lasso regularization zeroes out weak features).

In practice: scikit-learn

Use transformers and pipelines to apply feature steps consistently.

Real pipelines use scikit-learn's transformers, which fit on training data and apply consistently to new data:

from sklearn.preprocessing import OneHotEncoder, StandardScaler   # pip install scikit-learn
from sklearn.pipeline import Pipeline

scaler = StandardScaler()
X_scaled = scaler.fit_transform(X_train)     # learn mean/std from train
X_test_scaled = scaler.transform(X_test)     # apply SAME transform to test

scikit-learn snippet follows documented API

scikit-learn isn't installed here (the pure-Python transforms are run-verified). Its transformers are the standard tools — but note the critical pattern: fit on training data, then apply to test data. Computing the scaling from the whole dataset (including test) leaks information and inflates your results — a classic mistake.

Beware data leakage

The #1 feature-engineering pitfall: leakage — using information at training time that wouldn't be available at prediction time (or fitting scalers/encoders on test data). It makes your model look great in evaluation and fail in production. Always fit transforms on training data only.


Practice exercises

  1. Extend one_hot to handle unseen categories at prediction time (all-zeros row).
  2. Implement "robust scaling" using median and IQR instead of mean/std, and compare on data with an outlier.
  3. Create a price_per_unit feature from price and quantity columns, handling division by zero.
  4. Bin a list of ages into categories and one-hot encode the result.
  5. Explain data leakage with a concrete example and how fit/transform separation prevents it.

💬 Discussion

Have a question about this topic? Found an error? Share your thoughts below.