K-means: finding municipal profiles without inventing natural categories
Scaling, choosing k, silhouette, and stability for interpreting territorial clusters.
Series: from data to logistics decisions
- From orders to a logistics network: an optimization and machine learning project
- MILP: turning a logistics decision into a verifiable model
- Linear relaxation: how much better could a solution become?
- Greedy and local search: build quickly, then improve deliberately
- LNS: reorganizing part of a network to escape a local optimum
- Predicting demand: from a population baseline to Poisson and boosting
- Population forecasting: trends, damping, and temporal testing
- K-means: finding municipal profiles without inventing natural categories
- Candidate scoring: learning to filter without losing good decisions
- Scenarios and SAA: deciding before demand is known
Clustering answers a different question
Predicting orders estimates a quantity; clustering municipalities seeks similarities. K-means organizes observations around k centroids and reduces the sum of squared distances to those centers in feature space. Here that space includes population, growth, age, income, and distance to hubs. A centroid is a numerical profile, not a recommended warehouse location.
Think of organizing a library by similarity of subject matter: categories help exploration but depend on the chosen criteria. Changing features changes what counts as close. Without scaling, a variable with large numerical values can dominate another measured as a proportion. The module applies StandardScaler and multiple initializations to reduce dependence on an unfavorable starting point.

Prepare and run
import pandas as pd
from sklearn.preprocessing import StandardScaler
from alocacao_capacitada.ml.clustering import CLUSTER_FEATURES, choose_k, fit_clusters
table = pd.read_csv("results/estudo_integrado_20261001/dados_municipais.csv")
complete = table[table["completo"]].copy()
x = StandardScaler().fit_transform(complete[CLUSTER_FEATURES])
print(choose_k(x).to_string(index=False))
labels, profile = fit_clusters(table, k=4)
print(profile.to_string(index=False))The value k=4 demonstrates the interface; it is not a result-based recommendation. Read the candidate table first. Silhouette contrasts within-cluster closeness with separation from other clusters. The project’s stability procedure repeatedly fits on 80% subsamples, predicts labels for the full dataset, and compares partitions using adjusted Rand index. Seek interpretable groups that survive small sample changes.
From a label to a profile
Read municipality counts, median population, total population, income, and orders per group. More orders may simply reflect more people. Compare rates and dispersion as well as averages. Label numbers have no intrinsic meaning; the code reorders them by median population for readability, not as a quality ranking.
Ask whether the profiles help design scenarios or locate demand-model errors. If they merely produce attractive labels, a table by size and region may be enough. K-means favors compact groups under the chosen geometry and can obscure irregular structures or unusual municipalities. It does not establish causes: cluster membership alone does not explain why a city buys more.
Sources and evidence
Next: Candidate scoring: learning to filter without losing good decisions