A feature is one number or category that describes an example. A house listing becomes floor area, bedroom count, year built and postcode. An email becomes the number of links it contains, whether the sender is in your contacts, and which words appear. The model never sees the house or the email. It sees only these measurements, lined up as a feature vector, and learns how much weight to give each one.
That makes the choice of features a ceiling on what the model can learn. If the thing that really drives house prices, say the school catchment area, isn’t among the features, no algorithm can recover it.
Kinds of features
- Numeric: area, age, price. Usually rescaled with normalization so features with large values don’t swamp small ones.
- Categorical: colour, country, product type. These have to be turned into numbers, commonly with one-hot encoding.
- Derived: built from other features, such as price per square metre, or the day of the week taken from a timestamp.
Hand-made or learned
In classic machine learning, people design the features, and that work, feature engineering, often decides whether a project succeeds. Pedro Domingos put it bluntly in 2012: the features used are easily the most important factor, and building them is typically where most of a project’s effort goes.
Deep learning changed that for raw data like images, audio and text. A convolutional network is given raw pixels and learns its own features in its layers: edges, then textures, then object parts. This is representation learning, and it’s why “feature” also refers to the patterns a network learns internally. For tabular data, spreadsheets of customers and transactions, hand-built features still carry a lot of the weight.
The catch
More features isn’t automatically better. Each one adds a dimension, and the amount of data needed to cover the space grows very quickly; Domingos calls this the second biggest problem in machine learning after overfitting. Irrelevant features add noise, which is why feature selection exists.
Features can also leak the answer. A “refund issued” column predicts fraud almost perfectly in historical data, but at the moment you need the prediction, the refund hasn’t happened yet.