EstevezAlvarez
OptimizationPython

From orders to a logistics network: an optimization and machine learning project

What I built, how the algorithms connect, and what the capacitated facility location experiments demonstrate.

Series: from data to logistics decisions
  1. From orders to a logistics network: an optimization and machine learning project
  2. MILP: turning a logistics decision into a verifiable model
  3. Linear relaxation: how much better could a solution become?
  4. Greedy and local search: build quickly, then improve deliberately
  5. LNS: reorganizing part of a network to escape a local optimum
  6. Predicting demand: from a population baseline to Poisson and boosting
  7. Population forecasting: trends, damping, and temporal testing
  8. K-means: finding municipal profiles without inventing natural categories
  9. Candidate scoring: learning to filter without losing good decisions
  10. Scenarios and SAA: deciding before demand is known

The question before the algorithm

Where should distribution centers open, and which center should serve each region? Moving closer to customers reduces travel, but facilities cost money and each has limited capacity. My project turns that trade-off into a model for comparing decisions under explicit rules. This series teaches how to build and question that process, from the data to the interpretation of results.

DataPredictionDecisionEvaluation
Each stage answers a question: what we observe, what we expect, what we choose, and how we check the result.

A chain of distinct responsibilities

Data engineering preserves sources, harmonizes municipalities, and records provenance. Demand models estimate orders; population forecasts describe possible futures; clustering characterizes territorial profiles. Optimization then selects facilities and assignments. Scoring narrows the candidates, and scenario analysis checks how decisions respond to change. No stage replaces the next: predicting demand does not, by itself, determine where a facility should open.

Three scales of evidence

The teaching example has three centers and five regions: every decision can be enumerated, proving an optimum of 203 monetary units. The historical study uses Olist order subsets with 50 regions and 15 candidates, or 100 regions and 30 candidates. The national network adds population, propensity, and market scenarios. These experiments use different horizons and assumptions; their costs should neither be added together nor treated as directly comparable.

Saved results from the 2026-10-01 study; decimal points and model monetary units.
Method50×15 cost100×30 cost
Greedy994102.401352149.43
Local search920006.981339302.12
MILP919471.411310952.57
Warm-start MILP894531.261313883.90
LNS, mean of 3 seeds893976.561309923.03

The nominal budget was ten seconds per run; fast methods finished earlier, and LNS slightly exceeded the budget in some runs. All served 100% of the selected demand, representing approximately 27.66% and 43.01% of orders in the reference universe. Mean LNS cost was 10.07% below greedy cost on 50×15, using greedy cost as the denominator. This is a computational improvement under the stated assumptions, not a measured operational saving.

The project has three scales. The small example proves the optimum; larger studies compare methods; the national network explores scenarios.
The project has three scales. The small example proves the optimum; larger studies compare methods; the national network explores scenarios.

How to study and reproduce

Start with the model and run the small example before changing parameters. Then compare heuristics, bounds, and seeds; only afterward add prediction and scenarios. The repository requires Python 3.12 or later and declares Cavuca as an editable dependency in the sibling ../Cavuca directory: cloning this repository alone does not guarantee a complete installation. Series examples importing alocacao_capacitada require that prepared environment. The published website contains only static pages.

The local review on 2026-10-01 ran 91 tests successfully. It also identified a census metric that differs from conventional APE and an LNS description stronger than the code actually guarantees. The articles explain both limitations. Large-experiment results quoted here were checked against existing CSV files; not all national simulations were rerun. Real investment decisions still require freight contracts, operational capacity, complete costs, and service requirements.

Sources and evidence

Next: MILP: turning a logistics decision into a verifiable model

Back to the blog index