0% fanden dieses Dokument nützlich (0 Abstimmungen)
0 Ansichten206 Seiten

Solving Network Design

Die Dissertation von Andreas Bärmann behandelt Netzwerk-Design-Probleme in der kombinatorischen Optimierung, insbesondere im Kontext des Ausbaus der deutschen Eisenbahninfrastruktur bis 2030. Es werden verschiedene Lösungsansätze entwickelt, darunter Dekompositions- und Aggregationsalgorithmen sowie robuste Optimierungsmethoden, die auf reale Daten angewendet werden. Die Ergebnisse zeigen signifikante Verbesserungen in der Rechenzeit und der praktischen Anwendbarkeit der entwickelten Algorithmen.
Copyright
© All Rights Reserved
Wir nehmen die Rechte an Inhalten ernst. Wenn Sie vermuten, dass dies Ihr Inhalt ist, beanspruchen Sie ihn hier.
Verfügbare Formate
Als PDF, TXT herunterladen oder online auf Scribd lesen
0% fanden dieses Dokument nützlich (0 Abstimmungen)
0 Ansichten206 Seiten

Solving Network Design

Die Dissertation von Andreas Bärmann behandelt Netzwerk-Design-Probleme in der kombinatorischen Optimierung, insbesondere im Kontext des Ausbaus der deutschen Eisenbahninfrastruktur bis 2030. Es werden verschiedene Lösungsansätze entwickelt, darunter Dekompositions- und Aggregationsalgorithmen sowie robuste Optimierungsmethoden, die auf reale Daten angewendet werden. Die Ergebnisse zeigen signifikante Verbesserungen in der Rechenzeit und der praktischen Anwendbarkeit der entwickelten Algorithmen.
Copyright
© All Rights Reserved
Wir nehmen die Rechte an Inhalten ernst. Wenn Sie vermuten, dass dies Ihr Inhalt ist, beanspruchen Sie ihn hier.
Verfügbare Formate
Als PDF, TXT herunterladen oder online auf Scribd lesen

Andreas Bärmann

Solving Network Design


Problems via Decomposition,
Aggregation and Approximation
Solving Network Design Problems
via ­Decomposition, Aggregation and
Approximation
Andreas Bärmann

Solving Network
Design Problems via
Decomposition, Aggregation
and Approximation
With an Application to the Optimal ­Expansion
of Railway Infrastructure
Andreas Bärmann
Erlangen, Germany
[Link]@[Link]

Dissertation der Friedrich-Alexander-Universität Erlangen-Nürnberg, 2015

ISBN 978-3-658-13912-4 ISBN 978-3-658-13913-1 (eBook)


DOI 10.1007/978-3-658-13913-1

Library of Congress Control Number: 2016937515

Springer Spektrum
© Springer Fachmedien Wiesbaden 2016.
This work is subject to copyright. All rights are reserved by the Publisher, whether the whole or part of the
material is concerned, specifically the rights of translation, reprinting, reuse of illustrations, recitation,
broadcasting, reproduction on microfilms or in any other physical way, and transmission or information storage
and retrieval, electronic adaptation, computer software, or by similar or dissimilar methodology now known or
hereafter developed.
The use of general descriptive names, registered names, trademarks, service marks, etc. in this publication does
not imply, even in the absence of a specific statement, that such names are exempt from the relevant protective
laws and regulations and therefore free for general use.
The publisher, the authors and the editors are safe to assume that the advice and information in this book are
believed to be true and accurate at the date of publication. Neither the publisher nor the authors or the editors
give a warranty, express or implied, with respect to the material contained herein or for any errors or omissions
that may have been made.

Printed on acid-free paper

This Springer Spektrum imprint is published by Springer Nature


The registered company is Springer Fachmedien Wiesbaden GmbH
Danksagung

Vom optimalen Aufbau und Ausbau von Netzwerken handelt diese Arbeit, und in der Tat
hatte ich in den Jahren, in denen sie entstand, die Möglichkeit mir ein bemerkenswertes
Netzwerk von persönlichen Kontakten aufzubauen. Dafür, dass er mir diese Möglichkeit
gegeben hat, möchte ich vor allem meinem Doktorvater, Prof. Dr. Alexander Martin, dan-
ken, der von Anfang an meine Fähigkeiten geglaubt hat, und mich dabei unterstützt hat,
sie bestmöglich einzusetzen. Die Arbeitsatmosphäre an seinem Lehrstuhl ist geprägt von
Hilfsbereitschaft und einem sehr freundschaftlichen Umgang aller miteinander, der sehr
motivierend wirkt und es ermöglicht, sich mit jedweder fachlichen Frage oder Bitte an
seine Kollegen zu wenden  eine Sache von der ich sehr oft und gerne Gebrauch mache.

So ist es ganz natürlich, mich als nächstes bei meinen Kollegen zu bedanken, die mir
auf vielerlei Art und Weise geholfen haben, und denen ich hoentlich auch etwas davon
wiedergeben konnte. Zuvorderst danke ich meinem Bürogenossen Christoph Thurner, dem
besten den es gibt, für all die vielen Gespräche und Diskussionen, die wir hatten, und die
nicht nur das Arbeiten kurzweilig gemacht haben, sondern auch in bislang zwei gemeinsame
Veröentlichungen gemündet haben.

Auch meinen weiteren Koautoren möchte ich danken: Maximilian Merkert für das gemein-
same Knobeln an der besten Umsetzung für unseren Netzwerk-Aggregationsalgorithmus,
Dieter Weninger für seinen Einsatz, den entstandenen Code zu optimieren, und nicht zu-
letzt Prof. Dr. Frauke Liers, meiner Doktormutter, dafür, dass sie ihre groÿe Erfahrung mit
Netzwerkalgorithmen eingebracht hat, und dass sie uns wertvolle Tipps zum Aufschreiben
wissenschaftlicher Ergebnisse gegeben hat. Weiterhin danke ich Andreas Heidt für das ge-
meinsame Einarbeiten in die Welt der Robusten Optimierung und Prof. Dr. Sebastian
Pokutta für seine vielen Tipps und die Forschungsthemen, die er uns dabei erönet hat.

Einem Grenzgänger zwischen wissenschaftlicher Forschung und praktischer Umsetzung,


Hanno Schülldorf, möchte ich danken für die Expertise mit der er uns versorgt hat, als
es darum ging, die beste Modellierung für den Ausbau des deutschen Eisenbahnnetzes zu
nden, die Problemstellung, die die Motivation für diese Arbeit geliefert hat.

Unseren Rechnerbetreuern, Dr. Antonio Morsi, Dr. Björn Geiÿler, Denis Aÿmann und vor
allem Thorsten Gellermann danke ich für die Möglichkeit, jederzeit bei Ihnen vorbeizu-
kommen, und Hilfe bei Hard- und Softwarefragen zu bekommen. Bei dieser Gelegenheit
danke ich auch der Optimierungsgruppe von Prof. Dr. Michael Jünger an der Universität
zu Köln dafür, dass wir ihre Rechnerressourcen dort mitnutzen durften, und vor allem
Thomas Lange für seine technische Unterstützung dabei.

Meinem Kollegen Dr. Lars Schewe möchte ich dafür danken, dass er immer ein oenes
Ohr für komplizierte Optimierungs-Fragestellungen hat, und auf jede dieser Fragen eine
Antwort weiÿ.
VI Danksagung

Unseren Projektpartnern Prof. Dr. Ralf Borndöfer, Prof. Dr. Uwe Clausen, Prof. Dr. Ar-
min Fügenschuh, Prof. Dr. Christoph Helmberg und Prof. Dr. Uwe Zimmermann sowie
Dr. Boris Krostitz und ihren Mitarbeitern danke ich für den intensiven wissenschaftlichen
Austausch innerhalb unseres Forschungsverbundes KOSMOS, im Rahmen dessen diese Ar-
beit entstanden ist.

Oskar Schneider, meinem HiWi, danke ich für seinen Einsatz bei unserem neuen For-
schungsprojekt E-Motion, der es mir ermöglicht hat, mich mit mehr Zeit meiner Doktor-
arbeit zu widmen.

Prof. Dr. Martin Schmidt möchte ich dafür danken, dass er sich bereiterklärt hat, mein
2. Prüfer bei der Verteidigung zu sein; genauso danke ich auch Prof. Dr. Johannes Jahn
dafür, dass er dabei den Prüfungsvorsitz übernommen hat.

Sehr herzlich möchte ich unseren Sekretärinnen Christina Weber, Gabriele Bittner und
Beate Kirchner danken, den drei Damen vom Grill, die uns alle tatkräftig dabei unterstüt-
zen, unsere tägliche Forschungsarbeit durchzuführen.

Ich danke auch all meinen übrigen Kollegen am Lehrstuhl für die vielen Ausüge, Spie-
leabende, Theater- und Weihnachtsmarktbesuche und noch viele andere Dinge, die wir
gemeinsam erlebt haben. Mit euch zusammen hat es viel mehr Spaÿ gemacht, diese Arbeit
zu schreiben.

Zu guter Letzt danke ich meiner Familie und meinen Freunden. Durch euch konnte ich
diesen Weg überhaupt erst einschlagen.
Abstract

Network design problems are among the most-often studied problems in combinatorial
optimization. They posses a rich combinatorial structure that makes them interesting from
a theoretical point of view. On the other hand, they nd a huge variety of applications
in real-world problems settings, including transportation, supply chain management and
telecommunication  just to name a few. In this thesis, we develop solution approaches for
large-scale network design problems, focussing on dierent problem aspects.

Our motivation is a task set by our industry partner DB Mobility Logistics AG, namely
to come up with models and algorithms to determine an optimal capacity expansion of
line capacities in the German railway network until 2030. This aim is persued in Part I,
where we model the expansion of the railway network as a multi-period network design
problem. It is solved by a decomposition of the time periods along the planning horizon, for
which we give a heuristic and an exact variant. The algorithmic idea is easily transferable
to more general multi-period network design contexts. A dedicated case study for the
German railway network rounds this part o. It is based on real-world input data from
our industry partner and shows the practical relevance of our method. It produces very
satisfying results both from a mathematical and an applicational point of view.

In Part II, we focus on the spacial structure of network design problems. We develop an
exact algorithm for their solution that is based on aggregation of the underlying graph.
The idea is to solve the problem over a coarse representation of the original network and
to rene it iteratively where needed. We begin with a simplied version of the algorithm
for network design without routing costs. It takes the form of a primal cutting-plane
method whose cutting planes dominate the Benders feasibility cuts. Then we show how
routing costs can be integrated via lifted Benders optimality cuts. This leads to a hybrid
between aggregation and Benders decomposition and to a general idea how to realize a
Benders decomposition where variables are allowed to move from the subproblem to the
master problem. The aggregation scheme is tested on various benchmark sets in dierent
possible implementations. We are able to demonstrate signicant savings in computation
time compared to a solution via a standard solver.

Part III is dedicated to robust network design under ellipsoidal uncertainty. We develop
a framework for the approximate solution of the arising robust counterpart that is based
on a polyhedral approximation of the uncertainty sets. As an intermediate result, we
derive a new inner approximation of the second-order cone. The approximation framework,
which is in fact applicable to general mixed-integer programs, is again tested on various
benchmark sets. Our computational experiments show that the method is competitive
with traditional gradient-based linearizations and interior-point algorithms. Depending on
the type of problem to solve, it can be vastly superior to both of them.
VIII Abstract

Altogether, our work is a contribution to the practical solvability of network design prob-
lems, spanning the whole range from exact over approximative up to heuristic methods.
All the algorithms are tested on the network topologies provided by our industry partner
to demonstrate their applicability in real-world problem settings.
Zusammenfassung

Netzwerk-Design-Probleme gehören zu den meist-untersuchten Problemen in der kombina-


torischen Optimierung. Ihre reichhaltige kombinatorische Struktur und die groÿe Bandbrei-
te ihrer Anwendungen lassen sie sowohl aus theoretischer als auch aus praktischer Sicht
interessant werden. In dieser Arbeit entwickeln wir Lösungsansätze für groÿe Netzwerk-
Design-Probleme, wobei wir verschiedene Aspekte des Problems in den Fokus nehmen.

Unsere Motivation ist ein Forschungsauftrag unseres Industriepartners DB Mobility Logi-


stics AG: Die Entwicklung von Modellen und Algorithmen um einen optimalen Ausbau der
Streckenkapazitäten im deutschen Schienennetz bis zum Jahr 2030 zu bestimmen. Dieses
Ziel verfolgen wir in Teil I, wo wir die Erweiterung des Schienennetzes als mehrperiodiges
Netzwerk-Design-Problem modellieren. Wir lösen es mittels einer zeitlichen Dekomposition
entlang des Planungshorizontes, für die wir eine heuristische und eine exakte Variante ent-
wickeln. Die Grundidee lässt sich leicht auf allgemeinere Netzwerk-Design-Probleme über
mehrere Perioden übertragen. Eine ausführliche Fallstudie für das deutsche Schienennetz
rundet diesen Teil ab. Sie basiert auf realen Daten unseres Industriepartners und zeigt
die Relevanz unserer Methode für die praktische Planung. Die Ergebnisse sind sowohl aus
mathematischer Sicht als auch aus Sicht des Anwenders sehr zufriedenstellend.

In Teil II konzentrieren wir uns auf die räumliche Struktur von Netzwerk-Design-Pro-
blemen. Wir entwickeln einen Algorithmus zu ihrer Lösung mittels Aggregation des un-
terliegenden Graphen. Dabei wird das Problem zunächst für eine vergröberte Darstel-
lung des Originalnetzwerkes gelöst, und diese dann iterativ verfeinert. Wir zeigen, wie
sich eine vereinfachte Version des Algorithmus für Netzwerk-Design ohne Routingkosten
als primales Schnittebenenverfahren auassen lässt, dessen Schnittebenen die Benders-
Zulässigkeitsschnitte dominieren. Im zweiten Schritt integrieren wir die Routingkosten
mit Hilfe von gelifteten Benders-Optimalitätsschnitten. Das führt zu einem Hybrid zwi-
schen Aggregation und Benders-Dekomposition, bei dem Variablen vom Subproblem ins
Masterproblem wechseln dürfen. Wir testen das Aggregationsverfahren auf verschiedenen
Datensätzen und in verschiedenen Implementierungsvarianten. Es ermöglicht signikante
Einsparungen in der Lösungszeit verglichen mit den Ergebnissen eines Standardlösers.

Teil III widmet sich robustem Netzwerk-Design unter ellipsoidaler Unsicherheit. Wir entwi-
ckeln ein Verfahren zur approximativen Lösung des zugehörigen robusten Gegenstücks, das
auf einer polyedrischen Approximation der Unsicherheitsmengen basiert. Ein Zwischener-
gebnis dabei ist eine neue innere Approximation des Kegels zweiter Ordnung. Wir testen die
Approximationsmethode, welche tatsächlich für allgemeine gemischt-ganzzahlige Probleme
nutzbar ist, ebenfalls auf verschiedenen Datensätzen. Unsere Ergebnisse zeigen, dass sie es
mit traditionellen gradienten-basierten Linearisierungen und Innere-Punkte-Verfahren auf-
nehmen kann. Je nach Art des Problems kann sie beiden deutlich überlegen sein.
X Zusammenfassung

Insgesamt ist unsere Arbeit ein Beitrag zur praktischen Lösbarkeit von Netzwerk-Design-
Problemen, wobei wir die ganze Bandbreite von exakten über approximative bis hin zu
heuristischen Methoden abdecken. Wir testen alle Algorithmen auf den Netzwerktopologien
unseres Industriepartners, was uns erlaubt, ihre Anwendbarkeit für reale Problemstellungen
zu demonstrieren.
Contents

Danksagung V

Abstract VII

Zusammenfassung IX

Contents XI

List of Figures XIII

List of Tables XV

Introduction 1

I. Decomposition Algorithms for Multi-Period Network Design 7


1. Motivation 11

2. Strategic Infrastructure Planning in Railway Networks 15


2.1. Basic Terminology . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 15
2.2. Strategic Track Infrastructure Planning . . . . . . . . . . . . . . . . . . . . 18
2.3. The Problem in the Literature . . . . . . . . . . . . . . . . . . . . . . . . . . 28

3. Modelling the Expansion Problem 49


3.1. Modelling Assumptions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 50
3.2. Input Parameters . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 53
3.3. Single-Period Approach . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 56
3.4. Multi-Period Approach . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 59

4. Model Analysis and Solution Approaches 67


4.1. Comparing Models (FMNEP) and (BMNEP) . . . . . . . . . . . . . . . . . 67
4.2. A Compact Reformulation of Model (FMNEP) . . . . . . . . . . . . . . . . 70
4.3. Preprocessing of (NEP) and (CFMNEP) . . . . . . . . . . . . . . . . . . . . 72
4.4. Decomposition Algorithms . . . . . . . . . . . . . . . . . . . . . . . . . . . . 74

5. Case Study for the German Railway Network 85


5.1. Our Planning Software and the Computational Setup . . . . . . . . . . . . . 85
5.2. Evaluating the Models and their Enhancements . . . . . . . . . . . . . . . . 88
5.3. The Germany Case Study . . . . . . . . . . . . . . . . . . . . . . . . . . . . 90
XII Contents

II. Iterative Aggregation Procedures for Network Design Problems 103


6. Motivation 107

7. An Iterative Aggregation Algorithm for Optimal Network Design 111


7.1. Graph Aggregation and the Denition of the Master Problem . . . . . . . . 112
7.2. Possible Enhancements of the Algorithm . . . . . . . . . . . . . . . . . . . . 119

8. Integration of Routing Costs into the Aggregation Scheme 123


8.1. A Special Case: Aggregating Minimum-Cost s-t-Flow Problems . . . . . . . 124
8.2. An Aggregation Algorithm incorporating Routing Costs . . . . . . . . . . . 129

9. The Computational Impact of Aggregation 135


9.1. Computational Setup and Benchmark Instances . . . . . . . . . . . . . . . . 135
9.2. Computational Results on Scale-Free Networks . . . . . . . . . . . . . . . . 136
9.3. Computational Results on Real-World Railway Networks . . . . . . . . . . . 148

III. Approximate Second-Order Cone Robust Optimization 155


[Link] 159

[Link] Approximation of Second-Order Cone Robust Counterparts 163


11.1. Outer Approximation of the Second-Order Cone . . . . . . . . . . . . . . . . 163
11.2. Inner Approximation of the Second-Order Cone . . . . . . . . . . . . . . . . 166
11.3. Upper Bounds for Second-Order Cone Programs via Linear Approximation . 171
11.4. Approximating SOC Robust Counterparts . . . . . . . . . . . . . . . . . . . 172

[Link] Assessment of Approximate Robust Optimization 179


12.1. Implementation and Test Instances . . . . . . . . . . . . . . . . . . . . . . . 179
12.2. Results on Portfolio Instances from the Literature . . . . . . . . . . . . . . . 180
12.3. Results on Real-World Railway Instances . . . . . . . . . . . . . . . . . . . 185

Conclusions and Outlook 189

Bibliography 193
List of Figures

1.1. Development of rail freight trac in Germany between 2001 and 2013 . . . 11
1.2. Tracs ows under a 213 Gtkm scenario . . . . . . . . . . . . . . . . . . . . 13

2.1. The economic capacity of a link from the view of an RIC . . . . . . . . . . . 21


2.2. The economic capacity of a link from the view of an RTC . . . . . . . . . . 22
2.3. The economic capacity of a link for RTC and RIC together . . . . . . . . . 22
2.4. Capacity use along a sequence of links . . . . . . . . . . . . . . . . . . . . . 24
2.5. The dierent measures for increasing link capacity considered in our study . 28

5.1. GSV forecast for the growth in demand until 2030 . . . . . . . . . . . . . . 95


5.2. Our proposed expansion plan between 2010 and 2030 . . . . . . . . . . . . . 97
5.3. Chosen measures in our expansion plan for 2030 . . . . . . . . . . . . . . . . 99
5.4. The expansion plan for 2030 put forward by DB Netz AG . . . . . . . . . . 101

7.1. An aggregated graph . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 113


7.2. Illustration of the aggregation subproblem . . . . . . . . . . . . . . . . . . . 114
7.3. Disaggregation of an infeasible component . . . . . . . . . . . . . . . . . . . 115
7.4. Schematic outline of the three aggregation schemes . . . . . . . . . . . . . . 120

9.1. The aggregation algorithms on small scale-free networks . . . . . . . . . . . 138


9.2. Reduction of graph size for small scale-free networks . . . . . . . . . . . . . 140
9.3. The aggregation schemes on larger scale-free networks . . . . . . . . . . . . 141
9.4. Algorithms MIP and IAGG on larger scale-free networks . . . . . . . . . . 142
9.5. Reduction of graph size for some medium-sized scale-free networks . . . . . 143
9.6. Algorithms MIP and IAGG on small multi-commodity instances . . . . . . 145
9.7. Algorithms MIP and IAGG on medium-sized multi-commodity instances . 146
9.8. The aggregation algorithms on scale-free networks with routing costs . . . . 147
9.9. The aggregation algorithms on railway instances . . . . . . . . . . . . . . . 149
9.10. Algorithms MIP and IAGG on railway instances . . . . . . . . . . . . . . . 149
9.11. The aggregation algorithms on railway instances with routing costs . . . . . 151
9.12. Algorithms MIP and HAGGB on railway instances with routing costs . . . 151

11.1. The relationship between inner and outer m-gon approximations of B


2 . . . 166
11.2. Construction of the inner approximation of the unit disc B for k = 3
2 . . . 168

12.1. The quadratic algorithms on small portfolio instances . . . . . . . . . . . . . 182


12.2. The linear approximations on small portfolio instances . . . . . . . . . . . . 184
12.3. IP solver and linear approximation on medium-size portfolio instances . . . 185
12.4. IP solver and linear approximation on large portfolio instances . . . . . . . 185
List of Tables

2.1. Properties of the considered measures to eliminate bottlenecks . . . . . . . . 27

3.1. Summary of input parameters for our railway network expansion models . . 55

4.1. Data for the instance of Example 4.1.2 . . . . . . . . . . . . . . . . . . . . . 69


4.2. Results for the instance of Example 4.1.2 . . . . . . . . . . . . . . . . . . . . 70

5.1. Layout of the railway networks under consideration . . . . . . . . . . . . . . 87


5.2. Properties of the test instances for our railway network expansion models . 88
5.3. The railway network expansion models on the test instances . . . . . . . . . 89
5.4. The eect of preprocessing for Model (CFMNEP) . . . . . . . . . . . . . . . 91
5.5. The budgets allocated to each instance . . . . . . . . . . . . . . . . . . . . . 93
5.6. Performance of multiple-knapsack decomposition . . . . . . . . . . . . . . . 93

9.1. The aggregation algorithms on small scale-free networks . . . . . . . . . . . 137


9.2. Algorithms MIP and IAGG on small scale-free networks . . . . . . . . . . . 139
9.3. Selection rule for the larger scale-free networks . . . . . . . . . . . . . . . . 141
9.4. Algorithms MIP and IAGG on larger scale-free networks . . . . . . . . . . 142
9.5. Comparison of all algorithms on medium-size scale-free networks . . . . . . 144
9.6. Comparison of all algorithms on small multi-commodity instances . . . . . . 144
9.7. Comparison of all algorithms on medium-sized multi-commodity instances . 145
9.8. AlgorithmsMIP and HAGGB on scale-free networks with routing costs . . 148
9.9. AlgorithmsMIP and IAGG on railway instances . . . . . . . . . . . . . . . 150
9.10. Algorithms MIP and HAGGB on railway instances with routing costs . . . 153

12.1. Solution times on small portfolio instances . . . . . . . . . . . . . . . . . . . 181


12.2. Approximation accuracy on portfolio instances . . . . . . . . . . . . . . . . 183
12.3. Solution times on medium-size portfolio instances . . . . . . . . . . . . . . . 184
12.4. Solution times on large portfolio instances . . . . . . . . . . . . . . . . . . . 184
12.5. Tight approximation on railway instances with moderate uncertainty . . . . 186
12.6. Coarse approximation on railway instances with high uncertainty . . . . . . 188
Introduction

The optimal design of networks is an interesting combinatorial optimization problem which


has received a lot of attention in recent decades. In its simplest form, it may be stated as
follows. We are given a directed graph G = (V, A) with node set V and arc set A together
with a set R ⊆ {(v, w) ∈ V × V | v 6= w} of origin-destination pairs. Each of the arcs
a ∈ A incurs a cost ca if we choose to include it in the topology of the network. The task
is to choose a subset Ā ⊆ A of minimal cost such that the graph Ḡ = (V, Ā) contains a
directed path from v to w for each origin-destination pair (v, w) ∈ R. This is the so-called
uncapacitated network design problem in the case without routing costs, which means that
we pay no price for routing ow of any kind along the chosen arcs, but only for establishing
the point-to-point connections.

Even this simplest network design problem is very hard to solve in theory as it is a possi-
ble formulation of the NP-complete Steiner tree problem in graphs. In practice, network
design problems have already been used to model and solve many dierent problem set-
tings arising in real-world contexts. Applications include the design of transportation and
telecommunication networks, optimal chip layout in VLSI design, production planning in
supply chain management, the planning of distribution networks for water, electric energy
or natural gas or the determination of ancestor trees in genealogy, to only name a few.

To solve network design problems of real-world scale, it is usually necessary to devise


specialized solution algorithms, such as methods to decompose the problem into parts.
Together with modern branch-and-bound implementations, it is then possible to solve
them in an ecient manner.

Motivated by many practical applications, the optimal development of networks over time
is a trend which has gained more and more interest in recent years. This extension of
the network design problem adds a planning horizon T = {1, . . . , T } to the above setting
together with some kind of restriction on how many arcs may be chosen in each planning
period t∈T. The aim is to nd an optimal network development plan from a given initial
network in period 1 to some target network with the desired properties in the nal period T .
This class of problems, named multi-period network design problems or incremental network
design problems, is very well suited to model nancing plans for the gradual expansion of
existing networks.

A natural application of such a model is the expansion of infrastructure networks, especially


trac networks. Here, the construction of each new connection is very costly and the
annual budget of the network operator (very often the state) is limited. In this situation,
it is necessary to elaborate detailed plans for the expansion of the network over the planning
horizon as an ecient trac ow is not only aspired for the nal planning period, but also

© Springer Fachmedien Wiesbaden 2016


A. Bärmann, Solving Network Design Problems via Decomposition,
Aggregation and Approximation, DOI 10.1007/978-3-658-13913-1_1
2 Introduction

in any of the intermediate periods. This kind of problem can become very dicult to
handle as we have to incorporate |T | copies of the original network graph.

A further aspect is that the development of infrastructure to satisfy future demand scenar-
ios possesses a natural source of uncertainty as these demands can only be estimated, but
their exact realization is unknown at the time of planning. Thus, it has become common
practice to enhance network design models by incorporating techniques of stochastic opti-
mization and, more recently, robust optimization to account for this inherent uncertainty.
To forecast future demand patterns for the ow in networks, planners often use statisti-
cal evidence from past observations together with expectations for changes in the future.
In the context of stochastic optimization, this methodology suggests the consideration of
chance-constrained programs. A practical solution approach for this kind of problems
is their approximation via robust optimization models, which is possible by the suitable
choice of an ellipsoidal uncertainty set. The common formulation of the arising robust
counterpart converts the original mixed-integer program into a mixed-integer second-order
cone problem, which entails a signicant rise in complexity. As interior-point methods
possess striking disadvantages here due to the lack of warm-start capabilities within a
branch-and-bound procedure, the problem is usually solved via a gradient-based lineariza-
tion. However, this latter method suers from bad LP bounds until late in the procedure
when enough cuts have been added to represent the scenario set suitably well.

In the light of this discussion, this thesis sets out to nd new answers to some of the
questions brought up by the above developments:

How can we devise ecient algorithms to solve multi-period network design problems?
We will develop a novel method to decompose the problem along the timescale, allowing to
treat each time period individually. This is done exemplarily for the optimal expansion of
railway networks over time, which was the motivation for this thesis. We use the algorithm
to determine network development plans based on data provided by our industry partner
DB Mobility Logistics AG.

How can we eciently deal with the large-scale networks arising in real-world prob-
lems? We will devise an algorithm based on the aggregation of the spacial structure
of the network design problem, i.e. the clustering of individual nodes to components, to
drastically shrink the size of the underlying graph. We derive this algorithm for the case
without routing costs and show how it can be extended to incorporate routing costs via
the introduction of cutting planes. This method allows for signicant savings in computa-
tion time on the instances in our test set, which include networks from the above railway
optimization context.

How can the robust counterpart of a network design problem be solved eciently in
the common case of ellipsoidal uncertainty? We develop approximated formulations of
the second-order cone robust counterpart based on compact linear approximations of the
ellipsoidal uncertainty set. This includes the derivation of a compact inner approximation
of the second-order cone based on the decomposition scheme by Ben-Tal and Nemirovski.
Introduction 3

The computational results show that our approximation framework allows for considerable
reductions in solution time in many cases at few losses in solution quality. In addition,
the method also allows the use of coarser approximations to trade solution quality against
solution time. Finally, we show the suitability of the method to solve instances of a portfolio
optimization problem as well as original instances from railway network expansion, which
underlines the general character of the method.

Each of these three questions opens up a fascinating branch of research which promises
a better theoretical understanding of the problems under consideration and an increasing
range of solvable application problems at the same time.

The Structure of this Thesis


The present thesis is organized in three parts, each one dedicated to one of the topics raised
above. The outline is as follows.

Part I is about decomposition strategies for multi-period network design. It introduces the
problem of railway infrastructure development posed by our partner DB Mobility Logistics
AG, which was the starting point of our work. We explain the necessary basic knowledge
in infrastructure planning and give a classication of the problem in the context of network
design problems. In the following, we derive a simplied single-period model for the task
together with two equivalent models formulating the actual planning problem over multiple
time periods. After determining the theoretically better of these two multi-period models,
we devise several model enhancements to come to a more compact formulation. With
this point reached, we will use the given example of optimal railway network expansion
to work out an ecient heuristic decomposition scheme for multi-period network design
whose basic idea is transferable to problems beyond this special application context. It
is built on seperating the choice of the arcs in the target network from scheduling their
implementation and from routing the ow, making use of the introduced single-period
model as a master problem. Although not needed to obtain high-quality results, we also
satisfy the theoretical interest in an exact algorithm for the problem by deriving a suitable
extension of the heuristic which is guaranteed to converge to the optimal solution. We
continue by demonstrating the eciency of the developed models and algorithms using a
real-world data set on the current German rail freight network which includes the internal
demand forecast for 2030 from Deutsche Bahn AG. We will see that our models are a very
adequate description of reality and that our algorithm leads to very satisfying results from
the view of expert planners in the eld. Thus, we expect that the planning software which
we developed on the basis of our methods can serve our industry partner as a valuable tool
for evaluating dierent strategies of infrastructure development.

Part II introduces an algorithm for the solution of large-scale network design problems using
aggregation. It is built on the observation that for large underlying network graphs, the
optimal network design very often only requires a small fraction of the available arcs. This
observation is even more frequent when considering the optimal expansion of an existing
network. We develop an algorithmic framework which exploits this property by considering
aggregated versions of the network graph. The idea is to solve the network design problem
4 Introduction

over a coarser representation of the network and to rene this representation adaptively
where needed. This results in an aggregation algorithm with a master problem proposing
a network design for the coarse graph and a subproblem which checks the feasibility of the
proposal on the original graph. In the case of a feasible proposal, we are able to prove its
optimality for the original problem. In case of infeasibility, the subproblem delivers a set of
cuttings planes for the master problem to cut o this proposal. The introduction of these
cutting planes can be interpreted as a renement of the master problem network in certain
regions of the graph. We are able to show that they are theoretically stronger than the
Benders feasibility cuts arising from a related decomposition of the problem. The method is
rst presented for the simpler case of network design without routing costs. Later we show
how routing costs can be incorporated by the introduction of lifted Benders optimality
cuts. This yields a hybrid between aggregation and Benders decomposition which can be
generalized to a Benders decomposition which allows to shift variables from the subproblem
to the master problem in the process. Our computational experiments include results
for dierent possible implementations of the aggregation scheme. Altogether, they show
its superior performance compared to a solution of the problem via a standard solver.
The benchmark set includes network graphs from the railway network expansion problem
presented in Part I, considering dierent degrees of initial demand coverage.

In Part III of this thesis, we turn our attention to optimization under uncertainty  an
important aspect in many network design applications. More specically, we develop an
approximation framework for second-oder cone robust optimization which allows for a
faster solution of the corresponding conic-quadratic robust counterpart. Building on the
results of Ben-Tal and Nemirovski for a polyhedral outer approximation of the second-
order cone, our basic idea is a compact and tight polyhedral approximation of ellipsoidal
uncertainty sets. This will allow for a linear approximation of second-order cone robust
counterparts, which is especially desirable in the presence of integer decision variables. In-
stead of solving a conic-quadratic program with integer variables, the robust counterpart
stays an ordinary mixed-integer linear program. This enables an ecient warm-start of the
simplex-method at the nodes in the branch-and-bound tree. At the same time, working
with approximated ellipsoidal uncertainty sets allows to retain the underlying probabilistic
assumptions on the uncertain parameters, especially covariance information. This is a de-
nite advantage over polyhedral uncertainty sets as introduced by Bertsimas and Sim, which
lack this possibility. As an intermediate result of our work, we also derive a polyhedral
inner approximation of the second-order cone which is optimal with respect to the number
of variables and constraints. It complements the known outer approximation by Ben-Tal
and Nemirovski and attains the same approximation guarantees. In our approximation
of the robust counterpart under ellipsoidal uncertainty, the results on the quality of the
second-order cone approximation turn into error bounds for the security buer of each
robustied constraint. We test our approach on two dierent benchmark sets comprising
portfolio optimization problems from the literature as well as original nominal instances
from Part I on the expansion of railway networks. The results emphasize that our method
is suitable to obtain near-optimal robust solutions faster than by solution of the exact
quadratic model. Furthermore, it is able to produce high-quality solutions quickly by us-
ing coarser approximations. We will observe that much computation time can be saved in
comparison to standard mixed-integer second-order cone algorithms.
Introduction 5

At the end of the thesis, we will give our conclusions on the presented models and algorithms
as well as the theoretical insight gained in the course of our work. Besides solving the
originally posed problem to the satisfaction of our industry partner, we will see that the
question answered in the following chapters give rise to interesting new directions in the
practical solution of network design problems and beyond.

Incorporation of Joint Work with other Authors


This thesis incorporates collobarative work with other authors that has been published in
two joint papers.

Chapters 6 and 7 as well as parts of Chapter 9 are based on the article Solving Net-
work Design Problems via Iterative Aggregation, which is referenced here as Bärmann
et al. (2015b). It is joint work with Frauke Liers, Alexander Martin, Maximilian Merkert,
Christoph Thurner and Dieter Weninger and has appeared in Mathematical Programming
Computation. The algorithm for network design problems developed in this paper is aug-
mented by lifted Benders optimality cuts in this thesis such that it allows for the incorpo-
ration of routing costs on the arcs. Furthermore, we present new computational results on
instances from railway network expansion.

Part III is mainly based on the article Polyhedral Approximation of Ellipsoidal Uncer-
tainty Sets via Extended Formulations: A Computational Case Study, referenced here as
Bärmann et al. (2015a). This joint paper with Andreas Heidt, Alexander Martin, Sebas-
tian Pokutta and Christoph Thurner will appear in Computational Management Science.
The published results are again extended to network design computations from the context
of railway infrastructure development.
Part I.

Decomposition Algorithms for


Multi-Period Network Design
I. Decomposition Algorithms for Multi-Period Network Design 9

Multi-period network design has many practical applications. It is suitable for cases where
not only the desired target state of the network has to be planned, but also the evolution
of the network over time. The rst part of this thesis elaborates on an example of this
problem setting which arises in the context of planning railway infrastructure. It con-
tains the results achieved in a joint project with GSV, a divison of DB Mobility Logistics
AG which is concerned with the forecast and simulation of trac ows. Demands in rail
freight transportation are expected to grow signicantly over the next two decades, and
therefore they presented us with the question how they could plan the expansion of the net-
work capacities to accommodate these future demands using mathematical optimization.
A multi-period approach is mandatory here, as limited infrastructure budgets and the im-
plementation times of the infrastructure improvements require the provision of a detailed
expansion schedule. To nd an optimal evolution of the railway network with respect to
the protability of the transportation of the demand was our objective. The given task
was all the more interesting as planners at GSV itself were concerned with this problem
during the conduction of our research. This enables us to compare our own solutions to
those from experts in the eld. We will see that in this respect, our models and algorithms
lead to realistic assessments of network capacities and allow for viable proposals for their
extension.

Part I is structured as follows. In Chapter 1, we begin with a broader motivation of


the topic, giving more information about the recent and the predicted development of
rail freight trac in Germany. Chapter 2 introduces some basic terminology and conveys
fundamental knowledge of railway network planning as needed for the understanding of
our work. It continues with a classication of the considered problem into a hierarchy of
network design problems investigated in the literature, which is accompanied by a presen-
tation of the most common solution approaches. At the end, we distinguish our approach
to railway network design from previous approaches to the problem. With this background,
we are ready to derive mathematical models for the optimal expansion of railway networks
in Chapter 3. We rst develop a simplied planning model over a single time period,
denoted by (NEP), which can already be of high signicance to a planner as such. Even
more important is that this single-period model will be a basic ingredient in our algorithm
for the solution of the multi-period case. We derive two possible mixed-integer program-
ming formulations for this problem, named (FMNEP) and (BMNEP), for which we show
that (FMNEP) is the better one from a theoretical point of view. This is done in Chapter 4.
Furthermore, we derive a more compact reformulation of (FMNEP), which gives rise to
our nal model, denoted by (CFMNEP). For this model, we devise eective preprocessing
techniques which allow for a tremendous reduction in problem size by exploiting the special
structure of the objective function. Then we are ready to derive an ecent decomposition
method for (CFMNEP). It is based on dividing the solution of the problem into a multi-
step process named multiple-knapsack decomposition, in which each of the intermediate
results provides interesting information to a network planner. In the following, we show
how to extend this heuristic procedure to an exact method, by a suitable embedding into
a Benders decomposition scheme. The resulting algorithm allows to improve the solutions
of a slightly modied multiple-knapsack decomposition iteratively with guaranteed cover-
gence to the optimal solution. These procedures are derived using the example of railway
network expansion, but the basic considerations are general enough to allow the transfer
10 I. Decomposition Algorithms for Multi-Period Network Design

to other multi-period planning contexts as well. Finally, we demonstrate the eciency of


our methods in Chapter 5 by using them to solve the problem originally posed by GSV.
We conduct experiments on subnetworks of the German rail freight network to conrm the
superiority of Model (FMNEP) to (BMNEP) computationally and to show the signicant
impact of our model reduction to (CFMNEP) as well as our prepropecessing techniques.
Then we use our multiple-knapsack decomposition to solve the railway network expansion
problem on the Germany-wide instance provided by GSV, Our Germany case study ends
with a detailed discussion of the proposed infrastructure development plan and a compari-
son to the planning considerations of GSV itself. We will we see that we are able to provide
high-quality results which are convincing to experienced planners.
1. Motivation

The development of railway trac in Germany has shown a signicant increase over the
recent years, especially in rail freight trac. The total amount of goods transported on the
German railway network has risen from 291 Mt of transported goods in 2001 to 374 Mt
in 2013, which means a total increase by 29 % in 12 years or an average increase of 2.5 %
per year during this time. The growth was expected to be yet higher, and without the
eects of the economic crisis in 2009, the increase would probably have been even more
considerable. This development is shown in Figure 1.1. Experts in the eld speak of a mere
renaissance of rail freight trac, making the decline in earlier decades almost forgotten.

450
Megatonnes (Mt) of transported goods

400

350

300

250
1999 2002 2005 2008 2011 2014
Year

Figure 1.1.: Development of rail freight trac in Germany between 2001 and 2013 in
megatonnes (Mt) of transported goods per year; Statistisches Bundesamt (2014)

On the one hand, this development is explainable by an overall increase in freight trac,
which is due to the importance of Germany as an exporting country and also as a freight
transit country. Furthermore, ecological considerations gave rise to political incentives
aiming at the shift from road transport to rail transport. The latter plays a main role in
concepts and blueprints of planning authorities for a freight trac that is compatible with
the protection of the environment, particularly in terms of pollution, energy consumption

© Springer Fachmedien Wiesbaden 2016


A. Bärmann, Solving Network Design Problems via Decomposition,
Aggregation and Approximation, DOI 10.1007/978-3-658-13913-1_2
12 1. Motivation

and land use. In this respect, rail freight trac is widely seen as a resource preserving
means of transport. As a result, these incentives also led and still lead to remarkable
growth in inland railway trac.

A recent strategy paper by the Umweltbundesamt (2009), the German Federal Environ-
ment Agency, deems it possible to increase the modal share of rail freight trac from 18 %
in 2008 to 26 % in 2025. According to their gures, accomplishing this aim would require
a growth in the capacity of rail freight transport by 80 % over this period, aiming at a
value of 213 Gtkm in 2025. This would translate into 710 Mt of transported goods on the
railway network if we assume the mean transport distance of 300 km observed in recent
years to remain constant  a gure twice as high as nowadays.

The target value of 213 Gtkm for the transport capacity has been taken up in a study
commissioned by the Umweltbundesamt (2010) which investigates the expansion require-
ments of the German railway network to accomodate the resulting trac. According to this
study, the expansion of the railway network has dragged behind the demand development
for a long time as only the most pressing capacity shortages have been attacked and only
the most necessary investments in the sustainment of railway trac operations have been
undertaken. Furthermore, it diagnoses a bias in the previous investment strategy towards
projects focussed on passenger trac, neglecting the much higher growth in freight trac.
As a consequence, the German railway network exhibits several shortcomings in capacity
already in the status quo, which frequently causes disruptions in the operating schedule of
the railway transport companys.

The cited study aims to identify the regions in the railway network which need to be
extended most urgently within the next two decades. It takes the freight trac ows of
2007 with 105 Gtkm as a basis and projects that the current network without expansion is
able to accommodate between 10 to 15 % more trac (about 130 Gtkm). To predict the
trac ows under a scenario of 213 Gtkm per year, it uses the very simpliied assumption
that the volume of trac doubles on each of the lines in the network compared to the
values of 2007. Starting from these gures, it evaluates potentials for using the remaining
capacities to reroute part of the trac in order to identify the lines where bottlenecks are
expected to become most pressing. Figure 1.2 is directly taken from Umweltbundesamt
(2010) and gives a graphical representation of the bottleneck situations in their underlying
scenario.

The left picture shows that under their assumptions, signicant undercapacities in the
railway network have to be expected along the two main north-south corridors of trans-
portation in the west and in the centre of Germany as well as in other regions. On the right
picture, we see that using the remaining capacities in the network to reroute part of the
trac is not sucient to accommodate the complete demand without losses in operation
quality. Furthermore, it has to be taken into account that routing the trac along detours
decreases the eciency and thus the protability of the railway system.

From these ndings, the study comes to the conclusion that a determined expansion of the
capacities in rail freight trac is needed. They calculate with 11 billion Euros of required
capital to provide the necessary capacities and make suggestions for specic expansion
projects to enact.
1. Motivation 13

Figure 1.2.: Graphical representation of the tracs ows under a scenario of 213 Gtkm
of rail freight trac without (left) and with rerouting (right) to use remaining capacities
 routable trac in blue, remaining capacities in green, rerouted trac in yellow, missing
capacities in red ; source: Umweltbundesamt (2010)

The high growth rates in rail freight to be expected for the future together with the tight
public infrastructure budgets were the motivation for our industry partner GSV to start
a joint project with the FAU Erlangen-Nürnberg on optimal expansion strategies for the
German railway network. The aim of our joint study was to nd an optimal distribution
of the available budget to realize the infrastructure projects which are most benecial
for increasing the capacities in rail freight trac. This spacial distribution of the budget
was to be complemented by an optimal schedule for the implementation of the capacity
expansions. The ultimate goal was the development of a dedicated planning software to
be used by planners at GSV to study the optimal expansion of the network under dierent
demand scenarios.

Our partners at GSV provided us with the necessary data on the railway network, the
expected demand development and the eects of the available measures to improve the
infrastructure. Our study uses 2010 as its base year, for which we could use demand
gures from the database of GSV. For the target year 2030, they put especially high
eorts into realistic assessments of future demands, which are the foundation of any reliable
investment stategy. This demand prediction for the year 2030 serves as the basis for the
strategic planning of Deutsche Bahn AG as a whole.
14 1. Motivation

The resulting mean annual growth rates of about 2 % between 2010 and 2030 are not as
progressive as those underlying the study by the Umweltbundesamt (2010) with about 4%
over the same time horizon. Nevertheless, the GSV prediction results in a total increase
of almost 50 % in rail freight trac over 20 years, which can only be accommodated if the
throughput of the network is signicantly improved.

The above task leads to a challenging optimization problem whose solution is the motiva-
tion and the starting point of this thesis.
2. Strategic Infrastructure Planning in
Railway Networks

This chapter presents the most important concepts in railway infrastructure planning as
they are referred to throughout the thesis. After the introduction of some basic terminology,
we explain the necessary knowledge in railway infrastructure development. For most of
the technical descriptions, we lean on the work of Hörl (1998). Further details can be
found in Berndt (2001), Sewcyk (2004), Kettner (2005), Ross (2001) and Vieregg (1995).
In the last part of this chapter, we give a classication of the considered problem within
the context of the existing literature on network design problems in general and optimal
railway network planning in particular. This includes the establishment of a hierarchy of
network design problems for further reference within this thesis and a review of the most
popular solution methods.

2.1. Basic Terminology


We begin by introducing some basic vocabulary in the eld of railway transportation
as it will be used in this thesis. It focusses on the creation of optimal infrastructure
development plans for rail freight networks. Thus, our description of the structure of the
railway network and the organization of rail freight operations will be from the perspective
of freight trac.

2.1.1. Structure of the Rail Freight Network


The following paragraphs describe the main components of the rail freight network and
dene the technical terms to refer to trac in the network.

Stations In Ÿ 4 of the Eisenbahn-Bau- und Betriebsordnung (EBO (2012)), the German


Ordinance on the Construction and Operation of Railways, the term railway station is
dened as follows:

Railway stations are railway facilities consisting of at least one set of points
[`eine Weiche'], where trains are allowed to originate, end, give way or turn.

This basic denition already suces to understand how the term is used in this thesis.
We simply view them as the nodes in the network, which are connected by tracks, and
denote them by stations for short. The models developed here do not explicitly impose

© Springer Fachmedien Wiesbaden 2016


A. Bärmann, Solving Network Design Problems via Decomposition,
Aggregation and Approximation, DOI 10.1007/978-3-658-13913-1_3
16 2. Strategic Infrastructure Planning in Railway Networks

capacity restrictions on the stations. Furthermore, we consider the train ows calculated
by these models as broad estimations of capacity usage on the tracks. Therefore, we can
abstract from the processes of train blocking and reblocking, which are prevalent in the
operational planning of railway transport. Disregarding the classication process allows us
to ignore the dierent levels of hierarchy into which the stations are subdivided as well as
their classication capacities. For a more detailed treatment of the structure and functions
of dierent types of stations, we refer the reader to Berndt (2001).

Links, Tracks and Blocks The roadway which connects two stations in the railway net-
work and along which a train passes between the two is commonly called a line. Their
sections, connections between adjacent stations, form the edges of the network. Each line
consists of one or more parallel tracks and can be used in both directions. In case of an
even number of tracks on a line section, the tracks usually possess a xed assignment of
direction such that half of them is used for each direction. In case of an odd number of
tracks, one of them is usually operated in both directions. The latter requires crossing
points where a train driving in one direction can stop aside and let a train driving in the
opposite direction pass. In order to avoid confusion with the trains operated on a line ac-
cording to a given timetable, we use the term link instead of line section in the remainder
of this thesis.

For safety reasons, each link is divided into blocks. A train may only enter the subsequent
block on a link while it is not occupied by another train on the same track. The length of
the blocks on a link is called the block size and varies between the links in the network.
Two subsequent blocks are divided by signals which inform the train driver whether the
following block is occupied.

An important property of a link is the type of traction it allows for. Unelectried links
can only accomodate trains pulled by diesel locomotives, while electried links also per-
mit electric locomotives. The operation of the latter requires the presence of catenaries
and is more benecial as electric traction permits higher velocities and thus an increased
throughput at lower transportation costs in general.

Furthermore, the links in the railway network exhibit several more specic characteristics,
such as the permissible maximal velocity of trains using that link and the permissible types
of trains on that link, which are subsumed in their so-called line standard. Among these
characteristics, the capacity of a link is of especial importance to us. In short, we view
it as the value which states how many trains can pass the link on a given day in a given
direction. We will discuss this topic in some more detail in Section 2.2.2.

Note that there is an important dierence between a line as introduced above and a so-
called train path. While the former denotes the physical roadway for the trains, the latter
describes the availability of a track for a train run at a given point in time. The provision
of train paths is a central part of operational railway planning. With respect to the
infrastructure, it is closely connected with the train mix on a link. This term depicts the
succesion of trains of dierent types on a given link, which is an important factor for the
throughput and thereby the capacity it provides. However, it is practically impossible to
determine the train mix on a link for years in advance, for which reason it is futile to
2.1 Basic Terminology 17

incorporate it in a planning model of the scope of two decades as it is considered here.


Consequently, the provision of train paths and the resulting train mix are neglected in the
models developed in this thesis.

Train A train is a unit consisting of one or several locomotives pulling a certain number
of wagons or driving individually which runs from a given origin station to a given desti-
nation station within the railway network. The trains may either feature diesel or electric
locomotives, where the latter can only be operated on electried links.

Many links in the railway network are jointly used by both passenger and freight trac.
In our models, we will concentrate on an optimal allocation of capacity on the part on the
network used by freight trac together with an optimal routing of the latter. Therefore,
we will take the routing of the passenger trac on the jointly used links as a given base
load and consider only the remaining capacities for freight trac. Therefore, a train will
always be a freight train in the following.

Relation, Sample Day and Demand Scenario The term relation will be used in the
following to denote a triple of an origin station A, a destination station B and the total
number of trains which customers of the railway transport companies would like to send
from A to B on a given day. It therefore represents the cumulated demand between the two
stations. A sample day stands for the demand on a representative day within a planning
period (e.g. one year) and is given by the underlying set of relations. It is typical to
choose the demand level for this sample day to be around 80 % of the maximally expected
demand as 50 % are not enough to represent days with higher load, while 100 % would
be too conservative with respect to the network requirements. In the context of demand
prediction, we use the term demand scenario to discribe the collection of sample days
within the planning horizon.

The planning horizon of two decades, as considered here, does not allow for a more detailed
discrimination of the demand, e.g. by dierent kinds of goods. This is not only due to the
increase in problem complexity. It would be dicult to obtain the required data to begin
with. In this sense, it would be an overmodelling of the problem. The same applies
to the explicit consideration of train paths, as indicated above in the discussion on the
train mix. For similar reasons, our models do not consider the demand for each seperate
day in the planning horizon. Instead, they plan the allocation of capacities to enable an
ecient routing of the demand for each of the representative sample days, which is done
by identifying potential bottlenecks and the optimal way to resolve them. To allow for the
amortization of the investments undertaken late in the planning horizon, it is typical to
consider a subsequent observation horizon, in which they are only evaluated, but where no
further investments are considered.

2.1.2. Railway Companys


The Allgemeine Eisenbahngesetz (AEG (2013)), the German General Railway Law, intro-
duces the dierentiation between two kinds of railway companys. First, there are railway
18 2. Strategic Infrastructure Planning in Railway Networks

transport companies (RTCs), whose product are services concerning the transportation of
passengers and freight within the railway network. On the other hand, there are railway
infrastructure companies (RICs), which operate, construct and maintain the necessary in-
frastructure for the provision of these transportation services. The above-mentioned law
requires railway companys acting as both an RTC and an RIC to oer competing RTCs the
same conditions for the use of the infrastructure. This is to ensure a non-discriminatory ac-
cess to the railway network. The biggest RTCs in Germany are the subsidiaries of Deutsche
Bahn AG, although there are more than 300 smaller private RTCs. The dominant Ger-
man RIC is DB Netz AG, also a subsidiary of Deutsche Bahn AG, with negligible private
competition.

2.2. Strategic Track Infrastructure Planning


In our modelling of the expansion of railway networks, we take the point of view of a
(single) RIC which persues to increase the throughput of the network under its operation.
This is to be achieved via investments into the capacity of individual links subject to the
available budget. The objective in doing so is the design of an ecient network with respect
to the routing of the demand to maximize the protability of its operation by the RIC.

We begin begin by putting the considered planning problem into context with the dierent
scopes of infrastructure planning concerning the time horizon and the spatial resolution
of the network. This is followed by a presentation of the most important concepts to
analyse and describe the state of a railway network in terms of demand coverage and
operating conditions. Afterwards we give some technical background on potential measures
to increase the transport capacity of a railway link.

2.2.1. Planning Levels


The temporal scope of the infrastructure plan is the foremost consideration to be un-
dertaken in the treatment of the problem. In case of short- and medium-term planning
horizons (i.e. a few years), the RIC can mainly derive the requirements for the infrastruc-
ture from the status quo of the network together with the existing production plans of the
RTCs. The governing considerations on these two timescales are the disposal, inspection
and maintanance of the network. In contrast, the aim of long-term planning is the adaption
of the network topology to the demand patterns of the future, which includes the alleviation
of existing or expectable bottlenecks. For this purpose, the required demand information
has to be forecasted as the RTCs cannot commit themselves to the purchase of train paths
for many years in advance. Important gures in the estimation of future demands are the
extents of the expected trac ows and the competition with dierent means of transport,
among others. They are the basis for a prediction of the train quantities on each demand
relation.

Concerning the spacial scope, an integrated consideration of the complete network is of


utmost importance for an RIC in long-term infrastructure planning. It is the only way
to incorporate the interaction between trac volume, dierent production concepts of the
2.2 Strategic Track Infrastructure Planning 19

RTCs (in freight trac mainly block train transport, single-car transport and intermodal
transport, see Berndt (2001)) into the estimation of the resulting network load. As a con-
sequence, an optimal design of the whole network cannot simply be divised by assembling
development plans for individual subnetworks. Thus, we aim at the development of models
and algorithms which are capable of computing solutions for the German railway network
as a unity. This is partly in contrast to current planning practice, where the current Bun-
desverkehrswegeplan (BVWP (2003)), the German Federal Transport Infrastructure Plan,
is seen as a plan by call by its critiques. This term describes the perception that it
rather represents a collection of regional special interests than an integrated solution to
the requirements of the whole network.

Designing a network of very large scale together with the long time horizon is commonly
based on a so-called macroscopic network model. Such a model considers broad trac
ows as an abstraction of the movement of individual vehicles in a detailed network. The
opposite, a microscopic model includes a complete representation of local infrastruce such
as the platforms and points within a station as well as the detailed movements of the trains
within a station. Intermediate network resolutions are called mesoscopic models. A macro-
scopic network model is mostly created by aggregating the information of a microscopic
network, inter alia by aggregating the detailed representation of a station to a single node
in the network, neglecting train movements within the stations. The benet of this coarser
view is a decrease in the complexity of the data and shorter computation times for the
derived simulation and optimization models. It also reects the accuracy of the demand
data, which is based on estimates. The models for an optimal network design investigated
here will thus be based on a macroscopic infrastructure model. In order to accelerate the
solution of this type of problem, Part II features the development of aggregation algorithms
which are able to detect the interesting regions of a network automatically. This allows for
an even coarser resolution without losing relevant information.

Finally, a signicant amount of uncertainty is introduced into the problem by the long-term
demand forecasts. Thus, a planning horizon of this scope incurs the risk of laying out the
infrastructure for a demand pattern that will never arise. To prevent this from happening,
a common practice is to underly multiple demand scenarios which are determined by
dierent future trac expectations. Therefore, we will present approaches to account
for the uncertainty immanent in the problem by techniques of robust optimization in
Part III.

2.2.2. Link Capacities and Bottlenecks

In the present thesis, we aim to generate optimal expansion plans for existing railway
networks in contrast to designing new networks from scratch. Using an informal notion
of the terms capacity and bottleneck, an extension of link capacities is most essential
along lines on which the demand exceeds the ability to accommodate it. If certain links
along the line operate at their capacity limit and higher train numbers would be desirable,
they form trac bottlenecks. These two central terms to describe the state of a railway
network concerning its load under a given demand scenario are introduced in the following
paragraphs.
20 2. Strategic Infrastructure Planning in Railway Networks

The motivation of our work is to give decision support to the RIC operating the network to
devise a most protable expansion plan. In our macroeconomic approach, we will assume
that the RIC benets most if it acts such that the prots of its customers, the RTCs, are
maximized. This assumption is explained and justed as well.

Capacity An RTC's daily business is to transport goods (whose amount is given in tonnes
or number of trains) along a certain route (whose length is measured in kilometres). As
a result, it generates a certain trac volume (measured in tonnekilometres or trainkilo-
metres). For its use of the network infrastructure, an RTC pays fees to the infrastructure
manager, i.e. the RIC. In the following, we will take a viewpoint from which the allocation
of single transportation orders to complete trains with given origin and destination (the
so-called blocking process) has already been accomplished. This permits us to consider en-
tire trains as the quantity to be routed instead of single wagons. Furthermore, we neglect
the blocking capabilities of the stations as limiting factors for the trac on a link. This
allows for the following denitions.

Denition 2.2.1. The technical capacity of a link is the maximal number of trains that
can pass this link during a dened time interval (respecting the requirements imposed by its
technical limitations and ensuring a desired quality of service). In contrast, the economic
capacity of a link is the number of trains whose processing during a dened time inter-
val maximizes the protability of its operation (again respecting technical limitations and
quality of service).

According to the above denition, the technical capacity of a link is an upper limit on its
trac imposed by technical factors. It is principally determined by the minimal allowed
safety margins between two consecutive trains on that link. They are given by the minimal
headway time between two such trains. This security buer diers between dierent types
of trains involved. For example, it is signicantly longer for a high-speed passenger train
following a slower freight train than the other way round. Therefore, the train mix on a
link has great impact on its availability in daily operations. The consideration of mean
minimal headway times, as it is advisable in long-term planning, allows to abstract from
a concrete train mix. They can be determined by estimating the empirical probability of
any thinkable succession of trains on a link. Its throughput can then be calculated as the
quotient of the length of the reference period and the mean minimum headway time. By
increasing the mean minimal headway time by an adequate time buer to avoid congestions
and delays, we come to a practicable approximation of the technical capacity of a link.

On the other hand, considering the economic capacity of a link means taking a nancial
viewpoint, namely determining the optimal number of trains to maximize the prot of
the railway companies. It is determined by various parameters: its theoretic capacity, the
quality of the schedule, the quality of operations and the market conditions. The market
conditions comprise the temporal distribution of the trains over the day, the train mix,
delays, the required punctualities and specic revenues. These inuences contribute to the
respective revenue and expense curves of the links. For a more comprehensive exposition,
see Schwanhäuÿer (1995).
2.2 Strategic Track Infrastructure Planning 21

To assess the economic capacity of a link, we can plot the revenue and the expense curve in
dependence of the transported number of trains. This graph typically features a so-called
prots wedge, which refers to the part of the graph where the revenue curve is above
the expense curve. The schematic example in Figure 2.1 shows the determination of the
economic capacity of a link for the RIC managing it. The revenues of an RIC are the fees

Figure 2.1.: Assessment of the economic capacity of a link from the view of an RIC; adapted
from Schwanhäuÿer (1997)

charged from the RTCs for the use of the infrastructure, or more precisely, the leasing costs
for the train paths ordered by the RTCs. On the other hand, the RIC has xed costs for
buildings and facilities, the roadway, the power supply system and information systems as
well as variable costs depending on the trac volume. In the given example, we see that
its prot, given as the dierence between revenues and expenses, attains a maximum at a
trac of about 160 trains per day, given the limitations of the link under investigation.

This assessment is dierent for the user of the infrastructure. Figure 2.2 shows the prots
wedge of an RTC. Its income are the revenues for the transportation of goods from its
customers. They increase more or less linearly with an increasing number of transported
trains up to a point of saturation due to congestion eects on the links. Its expenses are
determined by the fees for route leasing, labour costs and time-dependent costs, where the
latter comprise the costs for the rolling stock and the wages of the sta involved in the
train run.

The picture indicates that there exists an optimal number of trains, 135 trains per day,
which maximizes the prot of the RTC on the considered link. Exceeding this value would
result in a signicant increase in expenses as congestion eects would lower the average
train speed. As a consequence, prots would fall or even turn into losses. Analogously,
transporting fewer trains than optimal results in prot loss by missed revenue.

We see that the optimal number of trains on a link can be dierent for both interactors. For
our study, we take a more integrated view of the railway system to come to a unied notion
of economic capacity. The route leasing fees can be interpreted as the contribution of the
RTCs to the investments of the RIC in the network in a system point of view. A dedicated
22 2. Strategic Infrastructure Planning in Railway Networks

Figure 2.2.: Assessment of the economic capacity of a link from the view of an RTC;
adapted from Schwanhäuÿer (1997)

study by Schwanhäuÿer (1997) shows that the sum of maximal prots of both the customers
and the operating company (i.e. the protability of the whole railway system) is attained
at the same number of trains on a link at which the RTCs attain their maximal prots.
This is represented in Figure 2.3, which is a superimposition of the previous two gures.
It features a net income area for the overall railway system, in which the summation of

Figure 2.3.: Joint assessment of the economic capacity of a link for RIC and RTC; adapted
from Schwanhäuÿer (1997)

the prots of RIC and RTCs as a function of capacity utilization can take positive values,
even if one the two operates at loss. Under the revenue and cost parameters chosen for
the diagrams, the net income area attains its maximum at 135 trains per day, which is
the same value that maximizes the combined prots of the RTCs. Note that the optimal
number of trains from the view of the RIC alone is higher. In case the two gures lie too
far apart, the RIC may act by increasing the fees for route leasing.
2.2 Strategic Track Infrastructure Planning 23

As a consequence of the above considerations, we base our models on the assumption


that the RIC takes its investment decisions such that the collective prots of both are
maximized in order to benet most in the long run. This is a well-established setup in the
literature on long-term planning of railway infrastructure (cf. for example Sewcyk (2004),
Kettner (2005), Hörl (1998) and Breimeier and Konanz (1994)). It has to be added that
the fees paid by the RTCs usually do not fully compensate the expenditures of an RIC.
For this reason, federal subsidies are necessary, which are justied by the public mandate
to advance and maintain the infrastructure.

The above gures insinuate that for an individual link there always exists a region with
positive prot. This does not necessarily have to be the case. A link can also operate
at a loss independent from the number of trains using it. However, such a link can still
contribute to the prot of the network as a whole. Furthermore, the optimal number of
trains to maximize the prots on one link does not necessarily coincide with the optimal
load on other links in the network. Considering the prot of a single link corresponds
to a local view on the railway network. Optimizing the number of trains to maximize
its prot individually may lead to a worsening of the prot of others. Thus, an optimal
load on one link might be suboptimal for the complete network. Consider the example
shown in Figure 2.4. It shows the situation on three successive links on a line between two
selected stations A and D with intermediate stations B and C. We see that the assumed
number of 180 trains per day going from A to D leads to an optimal use of capacity on
link CD, while the capacity is exceeded on links AB and BC  possibly due to dierent
track characteristics.

The above observation motivates transferring the notion of the economic capacity of an
individual link to that of the complete railway network to describe a load situation which
optimizes the global prot. Such a load situation coincides with an optimal routing of the
demand orders through the network, making use of the possibility to reject part of them
if their transportation is not protable. We will adopt this point of view in the models
presented in Chapter 3.

Bottleneck Via infrastructure measures, we can inuence the theoretic capacity of a


link, which, according to the above discussion, also inuences the economic capacity. Our
objective is to invest the annual infrastructure budget of the RIC to extend the theoretic
capacities in an optimal way. As the budget is limited, we need a criterion to determine
which links are most in need for extension. It is given in the following denition.

Denition 2.2.2. A link is called a quantity bottleneck if its economic capacity exceeds
its theoretic capacity.

Quantity bottlenecks are those links which restrict the protability of the railway op-
erations, as they do not allow to route the desirable number of trains due to technical
limitations. For our global view, we adapt the above denition by stating that a quantity
bottleneck is a link which limits the economic capacity of the network as a whole.

We remark that the literature also considers so-called quality bottlenecks (see e.g. Hörl
(1998)). This term describes a link on which the transportation costs are permanently
24 2. Strategic Infrastructure Planning in Railway Networks

Figure 2.4.: Illustration of capacity use along a sequence of links; adapted from Hörl (1998)
2.2 Strategic Track Infrastructure Planning 25

higher than the revenues (negative prot) or on which the prot is lower than for a com-
parable link. This kind of bottleneck is independent from the load of the link. It can
occur due to disadvantageous track propertie such as small curve radii or high inclina-
tions. Therefore, it can prevent new demands from arising due to the lower service quality
(e.g. by lower train speeds), which may limit the protability of the network, too. As the
track prole can hardly be changed by the measures considered in this study, we restrict
ourselves to the consideration of quantity bottlenecks, denoting them by bottlenecks for
simplicity in the following.

Mathematically speaking, bottlenecks form a cut in the underlying graph of the railway
network, which limits the network ow. In economic terms, they can be considered as
entities with singular marginal utility, i.e. increasing their theoretic capacity increases
the utility of the network, while increasing that of other links does not, as long as the
bottlenecks are not resolved. In both cases, a bottleneck leads to the non-satisfaction of
existing demands or their satisfaction at reduced prots or at a lower quality level. Thus,
they lead to missed prots by both operator and user of the network.

2.2.3. Investments in the Infrastructure


A link can become a bottleneck for several reasons that can be grouped into three areas:
infrastructure, rolling stock and logistics (see Hörl (1998)). Within each area, the RIC
has a dedicated set of measures to choose from in order to increase its theoretic capacity.
This thesis focusses on measures concerning the infrastructure, a term referring to immo-
bile facilities such as stations, tracks, signals, etc. The availability of an infrastructure
component is the fraction of time in which it can be used as planned and without any
restrictions. It has to meet the requirements security, reliability, sucient operation qual-
ity and eciency. In case of a bottleneck, not all of these requirements are met. Trac
loads beyond the theoretic capacity of a link can lead to disruptions in train ow in case of
unfavourable weather conditions, delays that are propagated through the network or other
operational disturbances. It may take a considerable amount of time for the aected part
of the network to recover and return to normal state. A bottleneck in the infrastructure
clearly results in a decreased quality of service.

The legal basis for investments into the Germany railway network is the Bundesschienen-
wegeausbaugesetz (BSchWAG (2006)), the Federal Railway Development Act. It mandates
the compilation of a requirements plan for the expansion of links in the railway network
which has to be in accordance with the Bundesverkehrswegeplan, the Federal Transport
Infrastructure Plan, especially when it comes to the planning for other means of transport
as well as combined trac. The requirements plan contains the infrastructure projects to
be realized and is updated every ve years in order to adapt it to developments in econ-
omy and transportation demand. Infrastructure investments comprising the construction,
extension and the replacement of links as well as the improvement of control systems are
mainly nanced by the Bund, the Federal Government. The costs for the maintenance and
repairs in the network are paid for by the RICs themselves.

In the following, we introduce the measures for the alleviation of infrastructure bottlenecks
considered in this thesis. It will be an important feature of the models developed later
26 2. Strategic Infrastructure Planning in Railway Networks

to give a monetary quantication of their eects with respect to the protability of the
network. For convenience, we denote the application of a given measure to a particular
link as an upgrade.

Construction of a New Link The construction of completely new links comes into con-
sideration if one of the following conditions is met:

ˆ There is no connection between two or more stations in the network but sucient
demand.

ˆ There is a connection, but the demand cannot be fullled due to limited theoretic
capacity.

ˆ There is a connection, but the demand exceeds the economic capacity, such that it
can only be fullled at increased costs for the operator

ˆ There is a connection with sucient theoretic and economic capacity, but either
high operating costs (e.g. due to an adverse track prole) or a low quality of service
(concerning, for example, the transportation time) make a new construction seem
protable.

New links may be constructed to replace existing lines in their full length or partly.

The construction of new links is the most powerful measure with respect to the elimination
of bottlenecks. At the same time it is the most expensive one (e.g. from 20 to 30 million
Euros per km for a double-track link for high-speed trac) and the most time-consuming
one (≥ 3 years) due to the long preparation, planning and approval phases. Furthermore,
it is the measure with the longest amortisation time and therefore requires a most careful
examination of the demand. Its benets not only lie in increasing the theoretic capac-
ity of the network, but also in creating shorter connections between two stations and in
opportunities for a homogenization of the train mix on other links.

Construction of New Tracks Increasing the number of tracks on an existing link is


favourable if its track prole allows for an ecient operation, but its theoretic capacity
is insucient. Besides the creation of new capacities, adding one or more tracks allows
for the reservation of tracks for the operation at certain speeds or directions and thus a
homogenization of the trac. Widening an existing link is in most cases easier to achieve
than an entirely new construction, both from the nancial view point (ca. 5 million Euros
per km and track) and the construction time (ca. 3 years). Therefore, it is a measure for
the medium- and long-term development of the network. We will restrict ourselves to the
extension of single-track links to double-track links and the extension of double-track links
to triple- or quadruple-track links.
2.2 Strategic Track Infrastructure Planning 27

Speed Improvements Increasing the admissible speed on a link can be achieved by a


variaty of single measures, like the elimination of level-crossings, the installation of faster
switches, the construction of passing tracks, equipping the link with a continuous automatic
train control, up to improving the track prole, e.g. by rectication. These measures
lead to higher and more homegeneous train speeds over longer distances, which is almost
indispensable on highly frequented links. Speed improvements vary in their cost (from 2
to 3 million Euros per km) and their implementation time (≤ 3 years) and are to be seen
as a short- to medium-term measure.

Electrication The electrication of a link comprises the erection of catenaries, facilities


for power supply and other necessary constructions to allow electric traction instead of
diesel traction. It is advisable if the permissible train weight on a link is too small using
diesel traction or if the transportation times are too long and cannot be shortened by other
means. Electrication is easy to implement on most links (costs of ca. 1 million Euros per
km, duration ca. 2 years) and is most ecient on links with high train density or with
long inclinations. It is therefore to be considered as a short-term measure.

Block Size Reduction The reduction of the block size of a link reduces the headway times
between two subsequent trains such that it permits a higher train density. It is preferable
on links where for historical reasons, the block size has not been dimensioned with respect
to an optimal capacitation but other criteria. Implementing this measure might also require
an installation of advanced control systems, such as continuous automatic train-running
control, to ensure the required level of security. Reducing the block size is the most
economical measure (ca. 1 million Euros per km) and is available within a short time
span (1 year). Thus, it is a short-term measure applicable as a quick remedy to alleviate
bottlenecks.

Summary Our case study for the German railway network presented in Chapter 5 includes
all the above measures except for the construction of new links. However, the incorporation
of the latter into the models and algorithms developed in the following is straightforward.
Table 2.1 summarizes the considered infrastructure measures and lists their properties as
we assume them for our computational results.

Infrastructure Cost per km Time for Corresponding


measure in [million e] implementation planning scope

New track 5 3 years medium- to long-term


Speed improvements 2 3 years short- to medium-term
Electrication 1 2 years short-term
Block size reduction 1 1 year short-term

Table 2.1.: Properties of the considered measures to eliminate bottlenecks

An illustration of the considered measures is shown in Figure 2.5.


28 2. Strategic Infrastructure Planning in Railway Networks

(a) The construction of new tracks is a pow- (b) An example of speed improvements is the
erful but costly measure to create new capac- construction of passing tracks. They allow
ity. Doubling the number of tracks roughly slower trains to give way to faster trains,
doubles the capacity of a link. Source: which increases the medium train speed.
Deutsche Bahn AG/Christian Bedeschinski Source: Deutsche Bahn AG/Günter Jazbec

(c) Electrication of a link allows electric (d) The reduction of the block size of a
trains to use it. They can drive at higher link allows for a denser succession of trains
speeds as diesel trains. Source: Deutsche by increasing the signal density. Source:
Bahn AG/Martin Busbach Deutsche Bahn AG/Jochen Schmidt

Figure 2.5.: The dierent measures for increasing link capacity considered in our study

2.3. The Problem in the Literature


The optimal design and expansion of railway networks has been an active topic of research
in recent years. From a mathematical point of view, it falls into the broad class of network
design problems (NDPs), which are an often-studied topic in the optimization literature.
They are both interesting from a theoretical and a practical point of view. On the one
hand, NDPs are a very general class of problems that can be specialized to many well-
known combinatorial optimization problems. The fact that an NDP can be used to model
the Steiner tree problem on graphs shows that it is NP-complete (see Garey and Johnson
(1979), Problem ND12). Further examples for its modelling power include the shortest-
path problem, the minimum-spanning-tree problem and the travelling-salesman problem.
On the other hand, NDPs nd vast application in real-world problem settings arising from
elds such as transportation, logistics and telecommunication, among many others.

The intention of this chapter is to give an overview of approaches from the literature for
2.3 The Problem in the Literature 29

the solution of NDPs in general and for railway network design in particular.

2.3.1. The Network Design Problem

Research on the network design problem dates back at least until the 1960s, were it was
investigated by Ridley (1968), Stairs (1968) and Scott (1969), among others, in the context
of transportation networks. Since then, numerous applications and algorithms to tackle
the problem have been examined, which is documented by the broad surveys given by
Magnanti and Wong (1984), Minoux (1989) and Balakrishnan et al. (1997). The unavail-
ability of more recent general surveys on the topic is surely not due to a decreased interest
in the problem, but rather due to the fact that its abundance in the literature warrants
more focussed reviews  such as that given in Costa (2005) on the application of Benders
decomposition for network design problems.

In the following, we give a short classication of the most important types of network
design problems with regard to this work. Furthermore, we summarize the most commonly
employed methods for their solution as a reference for the algorithms devised in this thesis.
For a more extensive presentation of approaches to the problem, we refer to the book
by Pióro and Medhi (2004), which focusses on the eld of telecommunication network
design.

A Hierarchy of Network Design Problems

Network design problems are a very versatile class of optimization problems. They can not
only be used to model many other prominent combinatorial optimization problems, but
they can also easily be adapted to many real-world application requirements. According
to the choice of variables and the types of side constraints that are added to the problem,
we can t the arising models into a hierarchy of NDPs. Each step deeper in the hierarchy
makes the NDP harder to treat from a computational point of view and thus requires more
sosticated methods to solve it. In the paragraphs below, we present a basic version of
such a hierarchy, which allows to classify the types of NDPs treated in the present thesis.

Uncapacitated Network Design The most basic setting of an NDP is as follows. We are
given a directed graph G = (V, A) together with a set R ⊆ {(v, w) ∈ V × V | v 6= w} of
node pairs representing the origin-destination pairs between which some kind of ow has
to be routed. The elements r∈R are frequently called commodities, a term that stands
for a kind of good to be transported through the network. Whenever we talk about our
application to railway network expansion, we will instead use the word demand relation
or relation for short. Furthermore, there are routing costs far for each unit of ow on arc
a ∈ A which may be specic to each commodity r ∈ R as well as costs ka for including
an arc a in the network design, i.e. setup costs for using it. The objective is to nd a
network conguration that minimizes the total cost incurred by setup and routing costs
together. This problem is called the uncapacitated (multi-commodity ow) network design
30 2. Strategic Infrastructure Planning in Railway Networks

problem (UNDP), a naming which refers to the fact that any chosen arc can be used to
transport an arbitrary amount of ow.

For the routing of the demand within G, we introduce continuous variables yar ∈ [0, 1]
for the fraction of commodity r ∈R routed along arc a ∈ A, which corresponds to the
assumption that the ow may split up arbitrarily at any node in the network. A second
set of variables ua ∈ {0, 1} models the choice of arcs to use for routing the ow. These
variables are modelled as binary as the full cost for any arc has to paid even when only
routing small amounts of ow along it.

This problem can be stated as the following mixed-integer program (MIP):

far yar
P P P
min ka ua +
a∈A r∈R a∈A

v = Or

1,  if
yar yar v = Dr
P P
s.t. − = −1, if (∀r ∈ R)(∀v ∈ V )
a∈δv+ a∈δv− 0,

otherwise

yar ≤ ua (∀r ∈ R)(∀a ∈ A)


u ∈ U
ua ∈ {0, 1} (∀a ∈ A)
yar ∈ [0, 1] (∀r ∈ R)(∀a ∈ A).

Its objective function is the sum of the setup and the routings costs. The rst set of
constraints are the so-calledow conservation constraints, which are used to model the
+
routing of the commodities. The notation δv stands for the subset of arcs from a ∈ A
− r
leaving node v ∈ V , while δv are the arcs entering this node. Furthermore, O and D
r

denote the origin and the destination of a commodity r ∈ R respectively. The constraints
state that the amount of ow of each commodity entering a given node has to equal the
amount leaving it, except for its origin and its destination where it has to depart and arrive
in full respectively. The second set of constraints is called linking or forcing constraints.
They imply the selection of an arc a if any ow is routed across it. Finally, constraints
u ∈ U oer the possibility to further restrict the topology of the network design. Such
restrictions could include multiple-choice constraints such as

X
ua ≤ 1
a∈A0

for a subset of arcs A0 ⊆ A or precendence constraints such as

u a1 ≤ u a2

for two arcs a1 , a2 ∈ A. It is also possible to consider a maximal budget B to be spent on


the network design, for example in the following way:

X
ka ua ≤ B.
a∈A
2.3 The Problem in the Literature 31

There are several straightforward variations of UNDPs, among which is the single-commo-
dity case, i.e. |R| = 1. Furthermore, the problem is often modied such that there are
multiple origins and destinations for each commodity, i.e. the ow conservation constraint
from above is replaced by:

X X
yar − yar = brv (∀r ∈ R)(∀v ∈ V ),
a∈δv+ a∈δv−

brv ∈ {−1, 0, 1} brv = 0.


P
with all and v∈V

Assuming far > 0 for all r∈R and a ∈ A, it is obvious that an optimal solution does not
contain cyclic ows. If this assumption does not hold, cyclic ows can be eliminated by
adding the following two constraints:

X
yar ≤ 1 (∀r ∈ R)

a∈δD r

and
X
yar ≤ 1 (∀r ∈ R)(∀i ∈ V \ {Dr }).
a∈δi+

An alternative way to model the capacity constraints is as follows:

X
yar ≤ |R|xa (∀a ∈ A).
r∈R

It could be suspected that this aggregate variant is more ecient as it involves a much
lower number of constraints. Actually, the contrary is the case. The disaggregate version
yields much better bounds in the LP relaxation, which reduces the number of nodes to be
explored during branch-and-bound.

Even this simplest network design problem is already NP-complete as it includes the Steiner
tree problem. Nevertheless, it possesses a favourable polyhedral structure, and several
authors have devised well-performing approaches to solve it. Hellstrand et al. (1992) were
able to show that the polytope describing the set of feasible solutions is quasi-integral,
which means that the set of its edges contains all the edges belonging to the convex
hull of the integer feasible solutions. Magnanti et al. (1986) solve UNDPs via Benders
decomposition, while Balakrishnan et al. (1989) develop a dual-ascent procedure for the
problem. In Holmberg and Hellstrand (1998), an ecent solution approach is developed
which is based on Lagrangean relaxation embedded into a branch-and-bound procedure.

Fixed-Charge Capacitated Network Design In the xed-charge capacitated version of


the network design problem (CNDP), we can pay a price ka to equip an arc a ∈ A with
a capacity of Ca . This may be done at most once  thus the name xed-charge. The aim
is to provide sucient arc capacities to route the desired demand through the network.
32 2. Strategic Infrastructure Planning in Railway Networks

Therefore, the demand of each commodity r∈R is now given explicitly by dr . A CNDP
can then by modelled as the following MIP:

far yar
P P P
min ka ua +
a∈A r∈R a∈A

v = Or

1,  if
yar yar v = Dr
P P
s.t. − = −1, if (∀r ∈ R)(∀v ∈ V )
a∈δv+ a∈δv− 0,

otherwise
(2.1)
dr yar ≤ Ca ua
P
(∀a ∈ A)
r∈R
u ∈ U
ua ∈ {0, 1} (∀a ∈ A)
yar ∈ [0, 1] (∀r ∈ R)(∀a ∈ A).

Variables yar ∈ [0, 1] and ua ∈ {0, 1} keep their respective interpretations from above, the
objective function and the ow conservarion constraints remain unchanged. The dierence
lies in the new capacity constraint, which now has to keep track of the amount of ow
routed along each arc. This ow on arc a∈A must be zero if the arc is not chosen and at
most Ca if it is chosen.

When passing from UNDPs to CNDPs, there is a signicant increase in computational com-
plexity. In computational practice, one often observes bad bounds from the LP relaxation
in addition to the high degeneracy already present in the uncapacitated case. Gendron
et al. (1999) solve CNDPs via Lagrangean relaxation, investigating the bound quality when
relaxing dierent subsets of the constraints. Hewitt et al. (2010) develop a neighbourhood
search with auxiliary integer programs for the arc-based formulation that is complemented
by a lower bound obtained from a path-based formulation. The latter is strengthed with
cuts discovered during the neighbourhood search. Their computational results show that
the method yields high-quality solutions quickly. Cutting planes for path-based formula-
tions of the problem can be found in Balakrishnan (1987). Chouman et al. (2009) devise
a cutting-plane method based on the so-called cutset inequalities, which estimate the ca-
pacity needed on subsets of the arcs inducing a cut in the graph (see Atamtürk (2003) and
the references therein). Stallaert (2000) introduces further valid inequalities together with
routines for their separation.

Network Expansion Network expansion problems (NEPs) are a generalization of the


above CNDPs where there are some previoulsy existing capacities ca ≥ 0 on each arc
a∈A that may be used at no cost. Furthermore, there is a set Ba of available modules
for each arc a ∈ A. The implementation of such a module b ∈ Ba increases the capacity
of arc a by b , which can be done as often as desired. Let the set B of all
Cb at a price of kS
modules be dened as B := a∈A Ba . An NEP can be stated as an MIP in the following
2.3 The Problem in the Literature 33

way:

far yar
P P P
min kb ub +
b∈B r∈R a∈A

1, if v = Or


yar yar −1, if v = Dr
P P
s.t. − = (∀r ∈ R)(∀v ∈ V )
a∈δv+ a∈δv− 0, otherwise

dr yar
P P
≤ ca + C b ub (∀a ∈ A)
r∈R b∈Ba
u ∈ U
ub ∈ Z+ (∀a ∈ A)(∀b ∈ Ba )
yar ∈ [0, 1] (∀r ∈ R)(∀a ∈ A).

Variables yar ∈ [0, 1] again model the fraction of commodity r using arc a, while variables
ub ∈ Z + now represent the number of times module b is installed on the corresponding arc.
The objective function models the total cost resulting from the routing and the network
design as before, but now it takes into account the individual costs of the modules. The ow
conservation constraints remain unchanged, while the capacity constraints now account for
the initial capacities and the additional capacities by the dierent available modules on
each arc.

An often considered additional side constraint which can be subsumed under the constraint
u∈U is a limited availability of each type of module. Let hb be the maximal number of
times module b ∈ B can be installed on the corresponding arc. Then such a constraint
could be incorporated as

u b ≤ hb (∀a ∈ A)(∀b ∈ Ba ).

Probably the rst publication on an NEP is by Christodes and Brooker (1974). They
propose a branch-and-bound-like procedure to solve it. Magnanti and Mirchandani (1993)
show that the problem is strongly NP-hard, even if there are no ow costs, only a single
commodity and only two types of modules for each arc. Another indication of the increased
complexity of the problem is given by Sivaraman (2007). He investigates a variant of the
problem where each commodity may only be split onto two dierent paths and shows that
this problem is APX-complete. Gendron et al. (1999) give a broad review on dierent
formulations and solutions methods for NEPs. Bienstock and Günlük (1996) study the
polyhedral structure of the problem and develop a cutting-plane algorithm based on facet-
dening inequalities which is able to solve real-world instances arising in telecommunication
network expansion satisfactorily. Frangioni and Gendron (2009) combine the use of cutting
planes and column generation for a binary reformulation of the problem.

Multi-Period Network Expansion The last stage in the hierarchy considered here are
multi-period network expansion problems (MNEPs). Here we are given a demand value
drt for each commodity r ∈ R which is dierent for each time period t in the planning
horizon T := {1, . . . , T }. Parameter T is the number of periods in the planning horizon.
34 2. Strategic Infrastructure Planning in Railway Networks

New capacities may now be installed in any period t∈T. A corresponding MIP looks as
follows:

kbt utb + far yart


P P P P P
min
t∈T b∈B t∈T r∈R a∈A

1, if v = Or


yart − yart = −1, if v = Dr
P P
s.t. (∀t ∈ T )(∀r ∈ R)(∀v ∈ V )
a∈δv+ a∈δv− 0, otherwise

drt yart Cb utb


P P
≤ ca + (∀t ∈ T )(∀a ∈ A)
r∈R b∈Ba
u ∈ U
utb ∈ Z+ (∀t ∈ T )(∀b ∈ B)
yart ∈ [0, 1] (∀r ∈ R)(∀t ∈ T )(∀a ∈ A).

The dierence to an NEP is that all variables, the objective function, the ow conservation
constraints and the capacity constraints have to be expanded over time. In the problem
variant stated above, additional capacities on any of the arcs have to be bought in every
period in which they are used. In this thesis, we will condsider a case where this cost is to
be paid only once (but stretched over multiple periods). This setting (but other modelling
considerations as well) calls for the inclusion of a monotonicity constraint for the network
expansion. If the additional capacacity on an arc is made available in one period, then it
should be part of the network design in any subsequent period as well. Such a constraint
could take the following form:

utb ≤ ut+1
b (∀t ∈ T \ {T })(∀b ∈ B).

Minoux (1987) investigates the theoretical and computational implications of passing from
an NEP to an MNEP and proposes a decomposition heuristic for the latter problem.
Bienstock et al. (2006) develop a model for an MNEP in the context of optical networks
that features continuous capacity upgrades, which is explained with the long planning
horizon and the assumption that scalable upgrade technologies will arise over time. They
consider the price of network usage as a variable which is coupled to the demand via a non-
linear relationship and incorporate a protection against link failure. The problem is solved
by projecting out the continuous capacity variables, yielding an algorithm that scales well
with the size of the instances. Pickavet and Demeester (1999) also consider a survivable-
network-design variant of the problem. They study the cost savings which are possible by
an integrated solution of the multi-period model compared to sequential upgrades that are
determined individually and come to the conclusion that the eect is indeed considerable.
Kalinowski et al. (2015), Baxter et al. (2014) and Engel et al. (2013) derive models and
algorithmic approaches for multi-period variants of dierent combinatorial optimization
problems which are special cases of multi-period network design.

Further Problem Variants Each of the above steps in the hierarchy of network design
problems may be further specialized by introducing new model requirements. The huge va-
riety of network design models investigated in the literature makes a complete enumeration
2.3 The Problem in the Literature 35

impossible. Thus, we only give a few examples of possible model extensions. van Hoesel
et al. (2004) present polyhedral results for non-bifurcated network design problems, i.e.
the case of unsplittable ow, as well as their bifurcated relaxations. Computational results
are presented for real-world telecommunication networks. A similar polyhedral study is
conducted in Atamtürk and Rajan (2002). Ljubi¢ et al. (2012) treat the case of a single
source and multiple destinations, while Chopra et al. (1998) focus on network design for
a single source and a single destination. The case of multi-period network design with
incremental routing is considered in Lardeux et al. (2007). In this variant of multi-period
network design, routing paths used in one time-period have to kept up in all later peri-
ods as well. Papadimitriou and Fortz (2014) present a model for time-dependent network
design, i.e. single-period network design with multiple demand matrices. The aim is to
minimize the combined network costs consisting of the costs for the chosen arcs and the
sum of the routing costs over the planning horizon. A vast amount of literature can be
found for network design under uncertainty. As an entry point for the interested reader,
we refer to Kerivin and Mahjoub (2005), Thapalia et al. (2012) and Raack (2012), which
treat survivable, stochastic and robust network design respectively.

Popular Solution Approaches for Network Design Problems

The available literature on network design problems introduces at least as many solution
procedures as there are models under investigation. In the following, we review the most
important solution approaches available, roughly categorizing them into the categories La-
grangean relaxation, Benders decomposition, column generation and further approaches.
We explain the basic functioning of these ideas at the example of the multi-period net-
work expansion problem (without additional side-constraints u ∈ U) and discuss their
advantages and disadvantages.

Lagrangean Relaxation The idea of Lagrangean relaxation is to reformulate complicating


constraints by relaxing them and penalizing their violation in the objective function. In
the case of a minimization problem, this approach yields a lower bound on the optimal
objective value which is at least as good as the LP relaxation of the problem. The ap-
proach is especially well suited when the resulting relaxed problem has a structure that
allows for the use of an eent specialized solution algorithm. On the other hand, the
constraints chosen for relaxation are very often ones that couple two or more subproblems
of the entire problem which could otherwise be solved seperately. The price to be paid
for this decomposition is (in most cases) a gap between the optimal value of the original
problem and the optimal value of the relaxed problem. Sometimes, it can be closed by
the introduction of cutting planes. To make Lagrangean relaxation an exact method, it is
most often necessary to embed it into an enumerative scheme such as branch-and-bound.
For a detailed introduction to Lagrangean relaxation, we refer to Lemaréchal (2001) and
Frangioni (2005).

In the case of network design problems, there are several possible choices for the con-
straints to be relaxed. A dedicated study on the resulting relaxations, their computational
behaviour and the quality of their bounds can be found in Gendron et al. (1999) for CNDPs.
36 2. Strategic Infrastructure Planning in Railway Networks

It turns out that the algorithm performs most favourable for the relaxation of the capacity
constraints. We sketch this approach in the following for MNEPs.

Lagrangean relaxation of the capacity constraints leads to the problem given henceforth:
!
kbt utb far yart λta drt yart Cb utb
P P P P P P P P P
min + + − ca −
t∈T b∈B t∈T r∈R a∈A t∈T a∈A r∈R b∈Ba

v = Or

 1, if
yart yart v = Dr
P P
s.t. − = −1, if (∀t ∈ T )(∀r ∈ R)(∀v ∈ V ) (2.2)
a∈δv+ a∈δv− 0,

otherwise

utb ∈ Z+ (∀t ∈ T )(∀b ∈ B)


yart ∈ [0, 1] (∀r ∈ R)(∀t ∈ T )(∀a ∈ A),

where λta ≥ 0 for t∈T and a∈A is the penalty parameter for violations of the corre-
sponding capacity constraint. The objective function minimized in Problem (2.2) can be
rewritten as:
XX XXX XX X
− λta ca + (far − λta drt )yart + (kbt − λta Cb )utb .
t∈T a∈A r∈R t∈T a∈A t∈T a∈A b∈Ba

Thus, for xed λ, it consists of a constant term, a term depending on the y -variables and
a term depending on the u-variables. The structure of Problem (2.2) allows to optimize
the last two terms seperately. More exactly, the problem decomposes into |T | · |R| single-
commodity minimum-cost ow problems and a binary problem solvable by inspection.

For each choice of λ, Problem (2.2) yields a valid lower bound for the optimal objective
value of the original MNEP. The aim of nding an assignment of the λ-values which yields
the best such lower bound leads to the so-called Lagrange dual, which is the optimization
problem given by
max L(λ)
(2.3)
s.t. λ ≥ 0,

where L(λ) denotes the optimal value of Problem (2.2) depending on the choice of λ.
The Lagrange dual (2.3) is a non-smooth optimization problem which is mostly solved via
subgradient or bundle methods.

The general advantage of this approach is the strong decomposition of the problem into each
time period and each commodity as single-commodity minimum-cost ows are easier to
compute than their multi-commodity counterparts. On the other hand, it can be expected
that even these single-commodity subproblems remain relatively hard  given the huge size
of network to be considered in this thesis, which lies in the order of 1600 nodes and 5000
arcs and for which about 3000 relations have to be routed in each of the 20 years of the
planning horizon. This is disadvantageous as these subproblems have to be solved very
often until the optimal penalty parameter λ has been determined. The strong coupling
of the networks under consideration contributes to this eect. Furthermore, the optimal
bound attained by the solution of the Lagrange dual is not better than the LP relaxation
2.3 The Problem in the Literature 37

of the original MNEP as Problem (2.2) possesses the integrality property (see Georion
(1974); cf. Gendron (1994) for a detailed proof for the case of CNDPs). To obtain an
optimal solution to the problem, it would be necessary to integrate the Lagrange dual into
a branch-and-bound framework, which would require even more subproblem evaluations.

Altogether, it seems that this approach does not scale for the problem under investiga-
tion here due to size of the underlying networks. In other problem settings, however,
Lagrangean approaches have succesfully been applied to network design problems. Gen-
dron et al. (1999) present an algorithm for CNDPs based on lower bounds determined
by Lagrangean relaxation of the capacity constrains. The Lagrangean dual is solved via
a subgradient procedure and the method is complemented by a heuristic to obtain fea-
sible solutions. They attain near optimal solutions for their benchmark set of random
instances. In a later paper, Gendron et al. (2001) use the same set of test instances to
show that the eciency of the algorithm can be increased in many cases by switching
from a subgradient method to a bundle method. Holmberg and Yuan (1998) introduce
a unied algorithmic framework based on Lagrangean relaxation for a variety of network
design models and show its eectiveness for large-scale instances. Sellmann et al. (2002)
develop a branch-and-bound algorithm where the bounds at each node are determined via
Lagrangean relaxation. The algorithm is enhanced by cardinality cuts and variable xing
rules. Chang and Gavish (1993) derive lower bounds for a multi-period network design
problem arising in telecommunications via a Lagrangean relaxation framework. In Chang
and Gavish (1995), they extend their work by employing a tighter model formulation.
Similar approaches for the multi-period case can be found in Kubat and Smith (2001) and
Dutta and Lim (1992).

Benders Decomposition The next solution approach under consideration is Benders de-
composition. It can be applied to linear and mixed-integer linear programs and relies on
the projection of the feasible set of the problem onto a subspace of the variables. By pro-
jecting out the continuous variables (or a subset of them), the total number of variables
can be decreased signicantly. This decrease comes at the price of a much higher number
of linear constraints which are necessary to describe the projection  in most cases expo-
nentially many. The algorithmic approach to deal with this problem is to generate these
constraints by demand, starting with a small (perhaps empty) initial subset of constraints.
The approach is developed for linear problems in Benders (1962). Generalizations of the
method to non-linear problems and to mixed-integer programming are given inGeorion
(1972) and McDaniel and Devine (1977) respectively. We give a sketch of the approach for
the case of MNEPs.

In our MIP formulation for the MNEP possesses two dierent types of variables: the inte-
gral variables u for the design of the network and the continuous variables y for describing
the ow in the network. The key observation here is that any xation of the u-variables
leaves a remaining problem in the y -variables which is much easier to solve as it decom-
poses into |T | independent minimum cost multi-commodity ow subproblems. For t∈T,
38 2. Strategic Infrastructure Planning in Railway Networks

they take the following form:

far yart
P P
min
r∈R a∈A

1, if v = Or


yart yart −1, if v = Dr
P P
s.t. − = (∀r ∈ R)(∀v ∈ V )
a∈δv+ a∈δv− 0, otherwise

drt yart Cb ūtb


P P
≤ ca + (∀a ∈ A)
r∈R b∈Ba
yart ∈ [0, 1] (∀r ∈ R)(∀a ∈ A),

where ū stands for the given xation of u. To obtain the cutting planes describing the
projection onto u, we need to consider the duals of these subproblems. For t∈T let αt
denote the dual variables to the ow conservation constraints, βt those to the capacity
constraints and γt those for the upper bound on y. Then the dual subproblem for each
t∈T reads:

rt − αrt ) − Cb ūtb )βat − γart


P P P P P
min (αO r Dr (ca +
r∈R a∈A b∈Ba r∈R a∈A

s.t. αvrt − αw
rt − drt β t − γ rt ≤ f r (∀r ∈ R)(∀a = (v, w) ∈ A)
a a a

βat ≥ 0 (∀a ∈ A)
γart ≥ 0 (∀r ∈ R)(∀a ∈ A).

Now, for each xation ū of u, there are two possibilities: the primal subproblems are either
all feasible or there is an infeasible subproblem for some t∈T. In the rst case, we need to
calculate the contribution of they -variables to the objective function. That means we have
to compute the optimal cost for a choice of y such that ū is complemented to a complete
solution to the MNEP as well as a lower bound for the y -cost in the vicinity of ū. This
objective function estimate φ for each t ∈ T comes in the form of a so-called optimality
t

cut, which is derived from the objective functions of the corresponding dual subproblem:
X X X XX
rt rt
(ᾱO r − ᾱD r − γ̄art ) − ca β̄at − (Cb β̄at )utb ≤ φt ,
r∈R a∈A a∈A a∈A b∈Ba

where (ᾱt , β̄ t , γ̄ t ) is an optimal dual solution for period t ∈ T . In the second case, the
xation ū makes one of the subproblems infeasible. Now, we know from LP duality theory
that there has to exist an unbounded ray in the feasible set of the dual subproblem. For
such an unbounded ray (ᾱt , β̄ t , γ̄ t ), we obtain a so-called feasibility cut :
X X X XX
rt rt
(ᾱO r − ᾱD r − γ̄art ) − ca β̄at − (Cb β̄at )utb ≤ 0.
r∈R a∈A a∈A a∈A b∈Ba

This feasibility cut can be interpreted as the minimal capacity to be installed on the arcs
belonging to some cut in the graph which would otherwise inhibit the routing of the desired
amount of ow in that time period.
2.3 The Problem in the Literature 39

The Benders master problem in the u-variables can now be written as follows:

kbt utb + φt
P P P
min
t∈T b∈B t∈T

rtp rtp (∀t ∈ T )


γartp ) − ca βatp − (Cb βatp )utb ≤ φt
P P P P P
s.t. (αO r − αD r −
r∈R a∈A a∈A a∈A b∈Ba (∀p ∈ P t )

rtq (∀t ∈ T )
rt − γartq ) ca βatq (Cb βatq )utb ≤ 0
P P P P P
(αO r − αD rq − −
r∈R a∈A a∈A a∈A b∈Ba (∀q ∈ Qt )
(∀t ∈ T )
utb ∈ Z+ (∀a ∈ A)
(∀b ∈ Ba ),
(2.4)

where (αtp , β tp , γ tp ) ∈ P t are the extreme points of the primal subproblem for t ∈ T
and (αtq , β tq , γ tq ) ∈ Qt the unbounded rays. Problem (2.4) is equivalent to the arc ow
formulation of the MNEP. It contains only the network design variables but possesses
an exponential number of constraints. The typical solution approach known as Benders
decomposition now starts with a restricted version of the master problem, i.e. with small
(or empty) subsets of extreme points and unbounded rays. In an alternating fashion, the
algorithm produces solutionsū to the restricted master problem which are used to nd new
optimality and feasibility cuts to cut o ū. As soon as no further such cut is found, the
algorithm terminates with an optimal partial solution ū to the MNEP. The accompanying
ows ȳ can be obtained by one more evaluation of the |T | primal subproblems.

Benders decomposition is widely used in the literature on network design problems as


the extensive survey by Costa (2005) shows. However, without exception, these applica-
tions focus on network design problems for comparably small underlying graphs which are
equipped with complex coupling constraints, such as it is the case in stochastic or robust
optimization, or in the presence of non-linear side-constraints. This is most probably due
to the often-observed numerical instability of Benders decomposition, which is caused by
the numerically dicult coecients in the Benders cutting planes. Furthermore, these
cutting planes are typically weak in the sense that many of them are needed to adequately
describe the feasible region of the master problem. Already for small network problems,
a standard text-book implementation of Benders decomposition is not sucient to obtain
an ecient algorithm. They require the derivation of stronger cutting planes which are
mostly problem specic. This is all the more true for bigger networks. To the best of the
author's knowledge, there are no prominent publications demonstrating an ecient use of
Benders decompositon for network design problems of the dimensions that are considered
in this thesis, which makes it very unlikely to be the algorithm of choice. Nevertheless,
we will see that the heuristic decomposition scheme developed in Section 4.4 exhibits a
close relationship to Benders decomposition. Therefore, we will be able demonstrate how,
in theory, this heuristic could be completed to an exact method by embedding it into a
specialized Benders decomposition.

Some selected applications of Benders decomposition to network design problems from


the literature are given henceforth. Magnanti et al. (1986) apply the method to solve
40 2. Strategic Infrastructure Planning in Railway Networks

uncapacitated network design problems. They show how cutting planes for this problem
that were previously known from the literature can be derived as Benders inequalities.
Furthermore, they introduce a procedure for a pareto-optimal lifting of the Benders cutting
planes that comes at the price of solving a number of minimum-cost network ow problems
equal to the number of commodities. It their computational experiments on randomly
generated instances, this lifting yield a considerable speed-up. Costa et al. (2009) classify
cutset inequalities as Benders inequalities and devise a method to strengthen Benders
inequalities to so-called metric inequalities. Examples for the use of Benders decomposition
for MNEPs are given by Dogan and Goetschalckx (1999) and Melo et al. (2005).

Column Generation Another often-applied approach for problems with an underlying


multi-commodity ow structure is column generation. The idea is to replace the pricing
step of the simplex algorithm by an optimization subproblem which identies the new
variable to enter the basis in each iteration. This may be useful when the problem possesses
a large number of variables as it avoids the consideration of the whole non-basis part of
the problem matrix to calculate the reduced costs. Therefore, the method is often used
in combination with a Dantzig-Wolfe reformulation of the original problem (see Dantzig
and Wolfe (1960)), which usually results in a model with fewer constraints but with an
exponential number of variables. Instead of introducing all these variables at once, new
columns are generated on demand. To obtain an exact algorithm, the approach has to be
incorporated into a branch-and-bound framework where it is necessary to allow for new
entering columns at all the nodes in the branch-and-bound tree  namely for those columns
needed for an optimal solution of the LP relaxation at that node. However, not all MIP
solvers support this feature. A very good introduction to column generation techniques
can be found in Desrosiers and Lübbecke (2005).

The Dantzig-Wolfe reformulation of a multi-commodity ow based problem typically in-


volves the use of path ow variables, i.e. variables describing the amount of ow on entire
paths in the network instead of ows on single arcs. For each commodity r ∈R let Pr
r r r r
P
denote the set of simple paths from O to D within G. Let Fp := a∈p fa be the total
r
cost of path p ∈ P for commodity r ∈ R, given as the accumulated cost for the arcs con-
stituting that path. Then such a reformulation of the arc ow formulation of the MNEP
might look as follows:

kbt utb + Fpr yprt


P P P P P
min
t∈T b∈B t∈T r∈R p∈P r

yprt = 1
P
s.t. (∀t ∈ T )(∀r ∈ R)
p∈P r

drt yart ≤ ca + Cb utb (∀t ∈ T )(∀a ∈ A) (2.5)


P P P
r∈R p∈P r : b∈Ba
a∈p
utb ∈ Z+ (∀t ∈ T )(∀a ∈ A)(∀b ∈ Ba )
yprt ∈ [0, 1] (∀r ∈ R)(∀t ∈ T )(∀p ∈ P r ).

Here, the y -variables have undergone a redenition. Variable yprt now stands for the fraction
of commodity r ∈ R which takes path p ∈ P r in period t ∈ T . We see that we can replace
2.3 The Problem in the Literature 41

the ow conservation constraints by (continuous) assignment constraints, which leads to


a reduction in the total number of constraints in the order of the number of nodes in the
network. On the other hand, we introduce an exponential number of path variables.

Using column generation, the problem can be solved as follows. Instead of the full set of
paths Pr for each commodity r ∈ R, we only consider a small (or empty) subset P̄ r . For
example, we could only consider the shortest path for each commodity, or the n shortest
paths for some n∈ N. This allows us to nd an initial solution for the branch-and-bound
root relaxation of the so-arising restricted Dantzig-Wolfe master problem. For an optimal
solution of the LP relaxation, we have to generate new columns with negative reduced
costs until no further such columns exist. The idea to nd such columns is to consider the
dual of the LP relaxation of Problem (2.5). It is given by

αrt − ca βat − γprt


P P P P P P P
max
t∈T r∈R t∈T a∈A t∈T r∈R p∈P r

αrt − drt βat − γprt ≤ Fpr (∀t ∈ T )(∀r ∈ R)(∀p ∈ P r )


P
s.t.
a∈A:
a∈p
Cb βat ≤ kbt (∀t ∈ T )(∀a ∈ A)
βat ≥ 0 (∀t ∈ T )(∀a ∈ A)
γprt ≥ 0 (∀r ∈ R)(∀t ∈ T )(∀p ∈ P r ),

where αrt are the dual variables for the assignment constraints, βat those for the capacity
rt
constraints and γp those for the uppers bound on y. For an optimal solution to the LP
relaxation of the restricted master problem, we consider the corresponding dual solution
(ᾱ, β̄, γ̄). We assume that there is a path p ∈ P r \ P̄ r for some commodity r∈R whose
inclusion would allow for a reduction of the optimal value. It follows from LP duality
theory that this is equivalent to the existence of a dual constraint

X
αrt − drt βat − γprt ≤ Fpr
a∈A:
a∈p

corresponding to this additional column which is violated by (ᾱ, β̄, γ̄) for some t ∈ T . How
do we nd such a violated constraint? This may be done by optimizing over the set of
simple paths for each commodity r and each period t in order to nd a path maximizing
the violation of the above constraint, given a xed dual solution (ᾱ, β̄, γ̄). As all drt are
non-negative, this amounts to the solution of |T | · |R| shortest-path problems to check
the dual feasibility of a solution (ū, ȳ) to the restricted relaxed master problem. Paths
corresponding to a dual infeasible solution are added to the restricted master problem,
and the process is repeated until dual feasibility has been proved. In this case, we know
that we have found an optimal solution to the LP relaxation of Problem (2.5).

Many heuristic schemes now work by switching back to an MIP by solving the restricted
master problem for the set of columns found so far. To come to an exact algorithm,
however, we have to check the integrality of the LP optimal solution; in case there is a
fractional u-variable, we have to proceed with a branch-and-bound scheme, repeating the
42 2. Strategic Infrastructure Planning in Railway Networks

above column generation for each of the node relaxations in the branch-and-bound tree.
This approach is called branch-and-price.

We note that the subproblems in the column generation procedure above again possess the
integrality property, thus the bound achieved by solving the LP relaxation of the Dantzig-
Wolfe master problem is not better than that obtained from the LP relaxation of the
original problem. Considering the specic structure of the problem solved in this thesis,
a further contraindication to column generation is that we will have to consider routing
paths from within a relatively large environment of the shortest path. That means there
will be a huge number of paths for each commodity which all lead to a similar cost in the
objective function and from which many similar paths may be active at the same time for
a given solution to the problem. Consequently, the restricted master problem will have to
incorporate a high number of similar columns at once, and the progress from generating
new columns can be expected to be relatively slow.

A more detailed introduction to column generation for network design problems can be
found in Ahuja et al. (1993). Frangioni and Gendron (2013) develop a stabilized column
generation framework with respect to the choice of the initial columns and to avoid the
tailing o eect, i.e. incremental progress when approaching the optimum. The algorithm
is applied to network expansion problems where it is shown to be eective. Gendron and
Larose (2014) develop a column generation method for xed-charge capacitated network
design that is based on an arc ow formulation similar to Problem (2.1). Examples for
the use of column generation for network design problems in practical applications are
Holmberg and Yuan (2003), Kim et al. (1999) and Cao et al. (2007)

Further Approaches There exists a huge variety of specialized methods for network design
problems, which is why we focus on the examples most relevant to our work on multi-
period network design. Apart from the generic approaches mentioned so far, mainly inexact
methods are used to solve this problem. These may broadly be classied into approximation
algorithms and heuristics.
The literature on approximation methods for network design problems is broad. For an
extensive survey, we refer the reader to Gupta and Könemann (2011). Kalinowski et al.
(2015), Baxter et al. (2014) and Engel et al. (2013) develop approximation algorithms and
heuristics for multi-period extension to the maximimum-ow problem, the shortest-path
problem and the minimum-spanning-tree problem respectively.

Garcia et al. (1998) develop a specialized local-search heuristic for an MNEP in the con-
text of telecommunication network planning. They show that it outperforms traditional
simulated annealing, tabu search and genetic algorithms. Kim et al. (2008) perform a sim-
iliar study, comparing three local-search algorithms for a multi-period problem modelling
highway network expansion. Gendreau et al. (2006) introduce an improvement heuristic
for the multi-period expansion of tree-shaped telecommunication networks that consists of
two phases. The rst one is a downstream pass from the leaves to the root to decide on the
installation of arc and node capacities. It involves the solution of an auxiliary multiple-
knapsack problem at each node in the tree to determine a least-cost expansion of this node
and its outgoing arcs with respect to a local price estimate. This estimate is updated in the
2.3 The Problem in the Literature 43

following upstream pass in reverse direction, and the procedure is iterated. They present
computational results for their method on tree networks with up to 110 nodes and tree
depth 9 over a planning horizon of 4 years. In Kouassi et al. (2009), the authors propose
a local-search heuristic and a genetic algorithm for the same problem, which enables them
to improve on their previous results.

A variety of methods developed in the literature work by reducing the solution of the multi-
period problem to the solution of its single-period version. In the short-term long-term
decomposition approach proposed in Minoux (1987), the demand of the nal time period
is used to determine a desirable target network as a solution to an NEP. In a separate
step, a feasible upgrade path from the initial network to the target network is determined
via a heuristic analogous to backward dynamic programming. In each time period, be-
ginning with the last one, the method tries to elimate the most expensive upgrades such
that the network design remains feasible for the previous period. Shulman and Vachani
(1993) extend the basic idea by iterating between the long-term and the short-term prob-
lem to improve the solution. Toriello et al. (2010) also considers a sequential decomposition
approach which is applied to an inventory routing problem. It does not prescribe a tar-
get conguration but incorporates approximate value functions to improve the obtained
decision path instead.

Our own algorithms for multi-period design presented in Chapter 4 are based on a short-
term long-term decomposition, too. Like the method by Minoux (1987), it starts by
deriving candidate upgrades from the demand pattern of nal time period. However, in our
approach, we do not try to nd a suitable upgrade path by considering only one period at
a time. Instead, we solve an auxiliary multi-period scheduling problem that incorporates
estimations on the protability of each upgrade in each time period. This way, we keep
a global view onto the whole planning horizon instead of a local view, which ultimately
allows us to obtain an exact method via an embedding into an enumerative scheme.

2.3.2. The Design and Expansion of Railway Networks

Network design and expansion models in railway transportation are a frequently studied
topic in the optimization literature. This is surely due to the tremendous capital required
to perform upgrades in the railway network over the typically long-term planning hori-
zons together with the otherwise almost unforeseeable consequences of combining dierent
upgrades. The approaches under investigation dier in the eects captured in the under-
lying model, the model formulation and the proposed solution algorithm. This section is
intended to present the publications most closely related to the content of this thesis and
to point out the new aspects incorporated here.

Literature Overview

The following literature overview is roughly arranged in decreasing order of relevance to


our own work. For each publication, we briey sketch the specic idea behind the chosen
44 2. Strategic Infrastructure Planning in Railway Networks

model, the algorithm employed to solve it and a summary of its results as far as they are
reported.

Lai and Shih (2013) consider a model for the upgrade of rail freight networks in North
America. The aim is a least-cost expansion of the existing capacities over multiple time
periods, taking into account upgrade costs, transportation costs and rejected demand.
In each period, the taken upgrade decisions have to respect the budget available for that
period. To incorporate a protection against uncertainty in the demand forecast, a stochastic
model extension is developed, which is solved via Benders decomposition. With their
approach, the authors are able to solve the stochastic model for an instance with 25 stations,
40 links and 20 relations over a planning horizon of 10 years clustered into 5 two-year
periods, considering 243 scenarios.

In Kuby et al. (2001), a single-period planning model is devised to solve the problem
of optimal network expansion for the Chinese railway system. This model is solved for
the last year of a given planning horizon and the set of chosen upgrades is then used
as input for the corresponding problem for the previous time period. This is continued
iteratively until reaching the rst time period to create a complete upgrade schedule. The
whole procedure is repeated several times with increasing underlying budgets for network
expansion to prioritize the chosen projects and to estimate their benets. The employed
model formulation features a mixture of path and arc ow variables, where the former are
used to restrict the routing to a set of predetermined paths in the network. Furthermore,
it allows for multiple upgrades on each link which may also be implemented in stages. Its
objective function minimizes the cost of the railway system consisting of routing costs,
construction costs and costs due to unsatised demand. Computational results for the
proposed algorithm are not reported.

Marín and Jaramillo (2008) investigate the problem of network expansion within the con-
text of urban rapid transit trac. They develop a model which simultaneously optimizes
the location of new stations, the choice of new links to construct and the oered origin-
destination transport services in the existing network. It is assumed that the demand
between two stations actually uses this public network if the transportation cost for the
user is lower than the cost of the competing modes of transport. The construction costs
are dened as the costs for installing a new transport service on one of the possible links
and have to keep to the budget given for each of the time periods. The model does not
consider an explicit capacity value for each link but assumes that whenever a link is con-
structed, the services operating on it can transport its entire demand. Their aim is to nd
a network evolution over multiple time periods such that an objective function combining
the demand coverage, the routing costs and the construction costs is optimized. For the so-
lution of large-scale instances, they propose a heuristic for the problem which decomposes
the problem into subproblems for each time period. Beginning with the data of the rst
time period, they compute an optimal single-period expansion of the network via branch-
and-bound. The optimal upgrade decisions are then taken as input for the single-period
expansion problem with respect to the second time period. This process is iterated until a
complete network evolution for the complete planning interval has been constructed. The
solution times and the solutions produced by this heuristic are compared to the results
of a standard branch-and-bound algorithm on three test networks. The biggest of them
2.3 The Problem in the Literature 45

models the network of the Spanish city of Sevilla and contains 24 stations, 264 links, 72
demand relations and up to 4 lines over 3 time periods. The heuristic is found to produce
satisfying solutions in much less time than needed via branch-and-bound.

Blanco et al. (2011) develop a multi-period planning model for the expansion of the Spanish
high-speed railway network. Its aim is to plan the construction of new stations and links
to cover a prescribed percentage of the country's population while fullling several other
requirements with respect to the level of service and the structure of the network. The ob-
jective function minimizes the sum of the construction costs over the planning horizon and
the routing costs arising over some observation horizon afterwards. The former addition-
ally have to respect a budget constraint for each period. Concerning the latter, dierent
economies of scale are assumed for the unit routing costs on single links and those arising
on paths with intermediate stations. Explicit link capacities are not taken into account.
The authors propose a scatter search heuristic to solve the problem. Starting with the rst
time period, it constructs a partial expansion plan up to the current period, which is then
passed to an improvement heuristic. The resulting solution is taken as input to construct
an expansion plan up to the subsequent time period, and the procedure is iterated until
it terminates with a complete plan for the whole planning horizon. The performance of
the heuristic is evaluated in comparison to an exact solution via branch-and-bound using
random problem instances with up to 100 stations and 30 time periods as well as a real-
world problem instance with 47 stations and 15 time periods. It is found to produce very
good solutions within much shorter time than the exact method. Moreover, the authors
show the validity of their model by demonstrating the close relation of its solutions to the
expansion strategy chosen by the Spanish government.

In Spönemann (2013), variants of single-period models for designing an optimal railway


network are introduced which are formulated using path ow variables and which allow
for dierent types of upgrades for the links. Furthermore, they allow for the consideration
of dierent train types and account for the train mix on a link. The models dier in the
capacity estimation resulting from the train mix. In both cases, the aim is to minimize
the construction costs necessary to enable a complete routing of the demand. The author
derives valid inequalities and preprocessing routines and proposes a column generation
method to solve the problem. Results are presented for dierent instances representing
Germany with up to 54 stations.

Petersen and Taylor (2001) consider the optimal design of a new railway line connecting
the north and south of Brazil to the agriculture and mineral resources of the northern
centre of Brazil. The line is planned over multiple time periods and maximizes the prot
of the railway operator, which is given by the dierence of revenues by accepting demand
orders and the costs arising from transportation and link construction. Their model does
not include capacities but assumes that any demand can be transported along a given link
once it is constructed. The size of the network under consideration is not exactly given,
but lies in the order of50 stations, 100 links, of which 10 are part of the line to be planned,
while the other ones refer to existing alternative modes of trac, 25 demand relations
and a planning horizon of 20 years. The authors develop a dynamic-programming scheme
which is able to solve the problem using a standard spreadsheet software.

In Repolho et al. (2013), the authors treat the optimal placement of stations along a
46 2. Strategic Infrastructure Planning in Railway Networks

new high-speed railway line within Portugals existing railway network. To this end, a
single-period model is derived to maximize the achieved savings in travel costs, of which
transportation time is taken into account as one important factor. The feasible set is
made up of set packing constraints to model the route choice of the passengers as well
as constraints to enforce a certain structure of the railway line. Their approach allows to
consider new demand generated by the network expansion. The results obtained from the
model are dicussed for the case of a new line between Lisboa and Porto, to which end it is
solved via branch-and-bound in conjunction with an ecient preprocessing routine.

Contribution of the Present Work

As indicated, our work is most closely related to that of Lai and Shih (2013). Their model
features the same ow conservation constraints to incorporate non-satisfaction of demand
orders as our models developed in Chapter 3. Furthermore, the formulation of the link
capacities is similar in the two approaches. While they assume that the capacity of a link
can be split arbitrarily between trac in both directions, we assume seperate capacities
for the two directions of a link.

On the other hand, our model introduces an important new feature concerning the realiza-
tion of the chosen upgrades. It explicitly takes into account that extensive measures in the
railway network usually take considerable time to implement. Therefore, their construc-
tion time may stretch over several periods of the planning horizon, while their eect, the
gain in capacity, is only available after the last period of construction. At the same time,
the costs for this investment have to be paid continuously during the time of construction.
This aspect is not present in any of the publications on multi-period network planning
cited above which all take the assumption that the whole investment cost is paid at once
after nishing construction. In this respect, our approach incorporates a ner modelling
of the nancial plan behind the network upgrade. In addition, construction times are also
important when considering investments that depend on each other. It may not be tech-
nically feasible to nish two measures in subsequent periods of the planning horizon if the
construction of the second one requires that the rst one has already been nished.

A further dierence to the work of Lai and Shih (2013) is that they assume that links can
only be upgraded once during the planning horizon, while they may be revisited in our
model. This represents a higher degree of exibility as links may be upgraded gradually
as needed. This aspect is treated in dierent ways in the above-mentioned publications.

Last but not least, our approach does not only introduce new modelling features, but we
also propose a new algorithmic way of handling the multi-period nature of the problem.
Many of the above publications propose heuristics reducing the multi-period problem to
the solution of one or several single-period versions of the problem. This also applies to
our approach. We will develop a decomposition heuristic that derives suitable candidate
upgrades from the solution to the single-period problem with respect to the demand pattern
of the nal time period, i.e. the target demand. Up to this point, this is a similar strategy to
the one employed in Kuby et al. (2001) or other publications. However, we do not try to nd
the best upgrade path to reach this nal choice of investments by unidirectional forward
or backward passes along the planning horizon. Instead, we propose a small auxiliary
2.3 The Problem in the Literature 47

scheduling model which incorporates estimations of the contribution of each upgrade to the
achievable increase in prot. These estimations are in turn calculated by solving continuous
minimum-cost ow problems for each time period. Altogether, our method takes a much
more global view onto the planning horizon than other methods, which certainly is one of
the reasons for its high-quality results, for which we refer to Chapter 5. Moreover, we will
show that our heuristic procedure can be extended to an exact method very easily.

The next steps will be to derive our modelling of the problem, performing model anal-
ysis and the derivation of the algorithmic scheme described above. We remark that this
derivation is given in terms of the application in railway network planning, but it is straight-
forward to extend our method to more general multi-period network design problems.
3. Modelling the Expansion Problem

The motivation behind the models developed in this chapter is the expansion of the current
German rail freight network to meet future demands. As outlined in Chapter 1, rail freight
trac is predicted to attain tremendous increases over the next two decades. On the other
hand, we saw in Chapter 2 that investments into the railway network bear a very high
price tag and need to be planned well in advance. Therefore, the question for an optimal
selection and scheduling of investments in accordance with the available budget is a vital
one.

We begin by summarizing and explaining the assumptions which guide the models pre-
sented here. Based on these asumptions, this chapter develops two modelling approaches
for an optimal expansion of the network which state the problem with respect to dierent
emphases. The rst approach considers the planning horizon as a single time period and
solely aims at determining a set of upgrades that is desirable to implement with respect
to developing the network. The second approach yields a temporally detailed expansion
plan for the network by considering the planning horizon as subdivided into multiple time
periods.

As indicated, the rst model focusses on the links to upgrade in order to optimize the
throughput of the network. It does not incorporate the scheduling of the upgrades and
only uses the demand situation of the last time period under consideration (the observation
horizon). The former means a simplication in terms of problem scale, which allows for a
quick evaluation of the available upgrades. The latter can be justied by the fact that the
nal demand situation represents the target trac for which the network is to be optimized.
We call this approach the single-period approach, as the underlying model only contains
one period for decisions (together with a second one for their evaluation). Consequently,
we assume that the total budget for the whole planning horizon is available at once. This
problem is formulated as a network expansion problem.

This single-period model will later act as an essential ingredient in an eective solution
algorithm for the multi-period approach. This ner modelling asks both for the upgrades
to be implemented and a schedule for their implementation. This has two adantages. It
allows for a more detailed analysis of the eects of the upgrades in the course the planning
horizon. On the other hand, it enables a ner view on their nancing. We develop two
multi-period network expansion models which mainly dier in the choice of the variables
representing the construction phase of an upgrade. In the rst model, it is incorporated
implicitly by introducing variables for the period of completion only. This leads to a more
complex statement of the budget constraint compared to the second model, which relies
on variables that explicitly state if an upgrade is under construction in a certain period.

© Springer Fachmedien Wiesbaden 2016


A. Bärmann, Solving Network Design Problems via Decomposition,
Aggregation and Approximation, DOI 10.1007/978-3-658-13913-1_4
50 3. Modelling the Expansion Problem

We will be able to show in Section 4.1 that the rst multi-period model dominates the sec-
ond one with respect to the strength of the linear programming relaxation. This theoretical
advantage within a branch-and-bound process will be conrmed by the computational re-
sults in Section 5.2.1.

3.1. Modelling Assumptions


In the following, we introduce the assumptions common to both the single- and the multi-
period approach before we derive the models themselves.

3.1.1. Link Capacities


The key question for upgrading the network is: Which links are bottlenecks or will become
bottlenecks in the course of the planning horizon due to rising demand? This question
mainly guides the investment strategy our models are to nd. Therefore, the most impor-
tant characteristic of a link is its throughput. As explained in Chapter 2, we distinguish
between the theoretic capacity of a link and its economic capacity. In short, the former is
the maximal number of trains which a link can accomodate, while the latter is the number
of trains which is most protable to operate on that link, both gures according to a cer-
tain time interval. Our models will mainly ask for an optimal investment in the theoretic
link capacities to allow for a most protable routing of the demand from the point of view
of the RTCs. This routing is part of the output of our models and can be viewed as a
certicate for the validity of the determined expansion plan. It can also be interpreted as
the determination of the economic link capacities.

As discussed in Chapter 2, we consider the theoretic capacity of a link independently from


its train mix, the order in which it is passed by dierent trains with potentially dierent
driving characteristics (speed, safety margins, etc. ). This corresponds to our of focus on a
rough long-term estimation of the trac ows in the network to determine the necessary
upgrades.

For convenience, we agree that whenever we use the term capacity without further at-
tribute, we mean the theoretic capacity of a link.

3.1.2. Passenger Trac


In Germany, rail freight trac and rail passenger trac mostly share a common network.
As already explained, our models view the passenger trac as a given base load for each
link which is subtracted from its capacity beforehand. Especially, passenger trac has a
xed routing determined a priori. Note that the remaining capacities might be dierent
for the two directions of a link because of this base load.

It is possible to model a change of the routing of passenger trac as a measure increasing


the capacity of certain links (where the passenger trac is taken away from) and decreasing
that of others (where the passenger trac is rerouted onto). That way, it suces to
3.1 Modelling Assumptions 51

consider a network with demands and capacities for the rail freight trac only. Accordingly,
measures only aecting passenger trac are neglected.

3.1.3. The Annual Budget


For the upgrade of the network, there is an annual budget available, which is mainly
provided by the state. It can be invested into the network according to a list of upgrades
which are available for each individual link. There are two possibile variants for modelling
the use of the budget. In the rst case, the budget obtained in any period can only be
spent in that particular period. Any remaining funds expire in that setting. The second
possibility contains the ability to set aside the remaining budget of each period for reuse
in later periods.

While we will present possible formulations for both alternatives, we will stick to a trans-
ferable budget for all further considerations. This is because there typically is no detailed
planning for the amount of money to spend in each period, but a rather a general idea of
how much will be invested in total. A transferable budget not only oers a higher degree
of exibility, but also allows to determine the amount of money that is actually needed
in each period. Starting from a rough estimation of the amount of money to be invested
per period, the reuse of any remaining budget may be understood as an indication that
less money is actually required in one period, while a later period requires more money
than initially allocated. Thus, the model does not only incorporate the decision into which
links the budget should be invested, but also how much money should be allocated to
each period. The latter is a valuable additional information which could very well be used
in budget negotiations to quantify the eects of the budget allocation over the planning
horizon.

In the single-period approach, we assume that we are given the complete amount of money
at once.

3.1.4. The Infrastructure Upgrades


We consider up to four dierent measures to increase the capacity of an existing link,
whose availability depends on the characteristics of the link: the construction of a new
track parallel to the existing ones, the electrication of a link, a velocity increase and a
reduction of the block size. The construction of a completely new link could be integrated
quite easily by modelling it as an upgrade for an existing link with an initial capacity
of zero. A detailed description of the measures is given in Section 2.2.3. Note that our
models do not explicitly include restrictions on the capacities of the stations. They could
be incorporated by splitting a station into two nodes which are joined by an an articial
link with limited capacity.

Each available upgrade, i.e. the application of a measure to a certain link, is dened by
three gures: its implementation cost, its (additive) contribution to the capacity of the link
and the time required for its construction. We take the assumption that the nancing of
each upgrade is equally distributed over the construction interval. Each upgrade increases
52 3. Modelling the Expansion Problem

the capacity of the link in both directions by the given amount. It could easily be rened to
individual payment plans per upgrade, maybe even representing dierent optional payment
plans for each project, if desired. The eect on the capacity of a link becomes available
after the last period of construction of an upgrade. Furthermore, we consider several
possible dependencies between the upgrades. On the one hand, there may be upgrades
whose implementation excludes those of others due to technical reasons. On the other
hand, there may be upgrades which require the completion of another upgrade because
they represent dierent stages of expansion within the same larger project.

In the single-period setting, we only determine the most protable combination of upgrades,
neglecting their implementation times. In the multi-period setting, our models accompany
the optimal choice of upgrades with an optimal expansion schedule.

3.1.5. Demand Orders


The criterion under which we optimize the expansion of the railway network is the achiev-
able prot by the transportation of freight demand orders. Therefore, our models in-
corporate the routing of the demand as a certicate for the optimality of the expansion
plan.

The long-term planning horizon justies the assumption that the routing paths of each
relation need not be unique, i.e. the ow may split at intermediate stations. Furthermore,
we assume that the blocking of the trains has already been planned, for which reason we
abstract from the types of goods transported, the train lengths, the number of wagons per
order, etc. Instead, the demand on a relation is given as the number of complete trains to
be routed, assuming an average load of 1000 t for each of them. Each order may be rejected
in total or partly if it cannot be routed in an economic fashion. This feature models the
incentive to choose those links for expansion which act as bottlenecks in accepting lucrative
orders.

The demand is assumed to be given for a sample day within each period of the planning
horizon. This sample day is chosen such that it is representative for each single day of
that period. To allow for the amortization of upgrades implemented late in the planning
horizon, we also incorporate a subsequent observation horizon, in which the demand is
assumed to stay constant.

3.1.6. Transportation Revenues and Expenses


The customer of an RTC typically pays a price for the transportation of his demand which
is proportional to the size of his order (here the number of trains) and to the length of the
shortest path in the network between its origin and destination. This is the revenue of the
RTC, if it accepts the order. On the other hand, it incurs costs proportional to the size of
the order and the length of the path along which it is actually transported.

It is immediately clear that this cost structure implies an interest of the RTC to ship
the orders along a shortest path to maximize its prot, i.e. revenues minus expenses. Any
3.2 Input Parameters 53

detour lets the prot shrink; an overly long detour, e.g. due to insucient capacities, might
even make it negative, in which case the RTC would rather reject the order. A partially
fullled order yields a revenue proportional to the fraction of trains transported. We will
see later that this assumption enables an ecient preprocessing of the models.

3.1.7. Objective
The objective of our models is an expansion strategy that leads to a maximization of the
prots achieved by the transportation of the demand orders. We do not distinguish between
dierent RTCs in our models such that fairness considerations are neglected. However, they
could well be included as possible extensions.

The costs for the expansion of the network are not part of the objective function as it is
typically assumed that the budget is completely used up. Instead, they are subject of a
budget constraint.

In the single-period approach, we only maximize the prot achieved within the observation
horizon. In the multi-period approach, we are to maximize the sum of prots in each
period of the planning horizon, during which upgrades may be implemented, and those of
the observation horizon, in which we only evaluate their eects.

3.2. Input Parameters


The models presented in this chapter mainly share the same input data, which is summa-
rized and explained here. It basically consists of four parts: the current state of the rail
freight network, (estimated) demand gures for each period, a list of available upgrades
for each link and the available budget for each period of the planning horizon. In the
following, we describe them in detail.

3.2.1. Planning Horizon and Observation Horizon


The planning horizon is the set of periods in which we can make investment decisions for
the network. We denote this set by T = {1, 2, . . . , T }, where 1 is the rst period of the
planning horizon and T T stands for the length of the planning horizon,
the last one. Thus,
too.

In each period t∈T, we receive a budget Bt, that can be invested into the network. As
discussed above, we state the modelling for both a reusable and a non-reusable budget but
will stick to the reusable case later.

The observation horizon is the set of periods following the planning horizon, in which
the eects of the implemented upgrades continue to be evaluated but in which no further
investment decisions are possible. As the demand is assumed to be constant over the
observation horizon, we can model it as a single time period after the planning horizon.
We thus denote it by the time index T + 1. The set T ∪ {T + 1} will be denoted by T̄ .
54 3. Modelling the Expansion Problem

Via the network ows over the periods of the set T̄ , the implemented upgrades are evaluated
in terms of their protability. The prots in time period T + 1 are weighted by W , which
represents the length of the observation horizon.

3.2.2. The Railway Network

The most important ingredient for our models is the railway network. As input, we are
given the initial state of the network at the beginning of the rst period of the planning
horizon. We represent this network as a directed Graph G = (V, A), where the node set
V A ⊆ {(v, w) ∈ V × V | v 6= w} is the set of tracks
is the set of stations and the arc set
connecting the stations. Especially, each link between two nodes v and w , denoted by
{v, w}, is modelled by the two arcs (v, w) and (w, v) representing its two directions. Note
that there may actually be several distinct tracks between a pair of nodes in each direction,
which the above notation neglects for the ease of exposition.

Associated with each track a∈A are the following parameters:

Length The length of a track in kilometres is denoted by la .

Capacity The initial capacity of a track is given by ca . It stands for the maximal number
of trains which can pass this track on each day. This could easily be extended to the
case of dierent base capacities in each period, which could be due to external eects.
The capacity of a track can be increased or decreased by upgrades undertaken during
the planning horizon.

3.2.3. The Demand

The relations are given in the form of a set of origin-destination pairs R ⊆ {(v, w) ∈ V ×V |
v 6= w}. Associated with each demand relation r∈R are the following parameters:

Origin and Destination For r = (v, w), we denote by Or = v and Dr = w the origin and
the destination of r respectively.

Train Count The size of a relation is given as the number of trains drt on a sample day
of period t ∈ T̄ which is to be transported from Or to Dr . It does not need to
be integral, but can also represent averaged demand values. Any relation can be
rejected partially or as a whole.

Distance The length of a shortest path from Or to Dr in G is denoted by Lr .

3.2.4. The Upgrades

The upgrades to increase the capacity of a track a ∈ A are represented by the set Ba . Each
upgrade b ∈ Ba is dened by the following parameters:
3.2 Input Parameters 55

Construction Time The realization of most upgrades cannot be completed within one
period alone. Therefore, mb is introduced for the number of construction periods
which b requires for its completion.

Implementation Cost The realization of upgrade b costs an amount of kb . In a multi-


period expansion plan, its payment is evenly split over the mb construction periods.

Capacity Increase The application of upgrade b to track a increases the capacity of the
latter by Cba . This notation allows for the case that an upgrade aects several tracks
at once.

For convenience, we introduce B := ∪a∈A Ba as the set of all upgrades available within the
network as a whole. Furthermore, we denote the subset of tracks aected by upgrade b∈B
by Ab .

3.2.5. Further Parameters


The remaining parameters are the following:

Revenue Factor The revenue of an RTC achieved by transporting one train along a dis-
tance of one kilometre is f1 . This value is an estimation for an assumed average load
of 1000 t per train.

Cost Factor The cost incurred by the RTC for transporting one train along a distance of
one kilometre is f2 . This value is also estimated for an average load of 1000 t per
train.

3.2.6. Summary of Parameters


Table 3.1 summarizes all the above parameters.

Parameter Unit Description

Bt [e] Budget in period t∈T


W [1] Length of the observation horizon (period T + 1)
la [km] Length of track a∈A
ca [1] Initial capacity of track a∈A
drt [1] Number of trains of relation r∈R in period t∈T
Lr [km] Length of a shortest path connecting Or ∈ V and Dr ∈ V
mb [1] Construction time of upgrade b∈B
kb [e] Total cost of upgrade b∈B
Cba [1] Eect of upgrade b ∈ Ba on the capacity of track a∈A
e
f1 [
km
] Revenue factor for accepting demand orders
e
f2 [
km
] Cost factor for transporting orders

Table 3.1.: Summary of input parameters for our railway network expansion models
56 3. Modelling the Expansion Problem

3.3. Single-Period Approach


In the single-period approach, the expansion of the railway network is modelled as a network
expansion problem (see Section 2.3.1) where the planning horizon is represented as a single
time step only. The available budget is equal to the sum of the budgets over the entire
planning horizon. The task is to nd an optimal selection of upgrades to implement with
respect to the prots attained over the observation horizon, taking into account possible
precedence relations as well as mutual exclusions among the upgrades. The certicate for
the validity of the chosen expansion plan is given by a feasible routing of the given demand
for a sample day of the observation horizon (period T + 1).

This problem description can be stated as the following optimization problem:

max Transport revenues - Transport costs

s.t. Route the accepted demand from origin to destination


Respect the track capacities
Keep to the available total budget
Respect the construction time of each upgrade
Respect the interdependencies between the upgrades

This problem as well as its algebraic concretion given in the following will be referred to
as (NEP) throughout the rest of this thesis. Note that we omit the time index T +1 for
parameters and variables for the ease of exposition.

The Variables

Our model for Problem (NEP) uses the variables introduced in the following. At rst,
there is a set of variablesub ∈ {0, 1} for each upgrade b ∈ B . It takes value 1, if upgrade b
r
is available in period T + 1 and 0 otherwise. Secondly, we introduce variables ya ∈ [0, 1] for
the routing of the demand. They stand for the fraction of demand relation r ∈ R which
r
is routed along arc a ∈ A. Finally, the model contains variables z ∈ [0, 1] for the fraction
r
of relation r ∈ R which is accepted. The value 1 − z can bee seen as the ow on an
articial track with innite capacity which directly connects the origin and the destination
of relation r.

The Objective Function

Problem (NEP) maximizes the prot made by the transportation of the demand. For each
relation r∈R the prot is modelled as follows. The customer pays a price of f1 per train
and kilometre along the shortest path, whose length is given by Lr . When accepting a
fraction of zr of the demand on relation r, the prot of the RTC is given by the term
f1 dr Lr z r . On the other hand, the expenses are proportional by f2 to the number of trains
and the distance actually travelled. Thus, the expenses for transporting relation r are
3.3 Single-Period Approach 57

f2 dr r
P
given by a∈A la ya . Summing over all relations yields the total prot of the RTCs in
each period of the observation horizon:

X X X
f1 dr Lr z r − f2 dr la yar . (3.1)
r∈R r∈R a∈A

The maximization of Function (3.1) is the objective of Model (NEP).

The Constraints

In the following, we derive the constraints subject to which Objective Function (3.1) is
optimized.

The Routing The requirement to route all accepted demands from origin to destination
can be modelled as a classical ow conservation constraint. The fraction of relation r∈R
emerging from Or is zr . The same fraction has to arrive at Dr . For the rest of the nodes,
inow has to equal outow. Thus, this constraint reads:

zr , v = Or

X X  if
yar − yar = −z r , if v = Dr (∀r ∈ R)(∀v ∈ V ). (3.2)
0,

a∈δv+ a∈δv− otherwise

The Link Capacities The total ow routed along any track of the network must respect
the available capacity, which can be increased via upgrades. The capacity constraint takes
the following form:
X X
dr yar ≤ ca + Cba ub (∀a ∈ A). (3.3)
r∈R b∈Ba

The Budget In every period t ∈ T , we receive a certain budget B t . From the


P pointt of
view of the observation horizon T + 1, we have obtained a total amount of t∈T B to
invest in the network. As the allocation of the budget to each single planning period is
neglected in this model, the choice between a transferable and a non-transferable budget
does not play a role, either. In both cases, Problem (NEP) is a relaxation of the multi-
period planning process with respect to the use of the budget.

We may not invest more money than obtained in total over the planning horizon. This
leads to the constraint:
X X
k b ub ≤ Bt. (3.4)
b∈B t∈T
58 3. Modelling the Expansion Problem

Interdependencies between the Upgrades The two cases of interdependencies between


upgrades considered here are mutual exclusion and required precedence. The rst case
can be modelled via multiple-choice constraints. Let Eb ⊂ B be the subset of upgrades
that is not compatible with a given upgrade b ∈ B. Then we have to respect the following
constraint:
X
ub + ub0 ≤ 1 (∀b ∈ B). (3.5)
b0 ∈Eb

The second case covers subsequent stages of a larger expansion project. Let Pb ⊂ B be
the set of upgrades which have to be completed before a given upgrade b ∈ B can be
implemented. Then we include the constraint:

ub ≤ ub0 (∀b ∈ B)(∀b0 ∈ Pb ). (3.6)

Complete Statement of (NEP)

Model (NEP) maximizes the RTCs' prot given by Objective Function (3.1) subject to
Flow Conservation Constraint (3.2), Capacity Constraint (3.3), Budget Constraint (3.4),
Multiple-Choice Constraint (3.5) and Precedence Constraint (3.6). Altogether, it reads as
follows:

dr Lr z r − f2 dr la yar
P P P
max f1
r∈R r∈R a∈A

z r , if v = Or


yar yar −z r , if v = Dr
P P
s.t. − = (∀r ∈ R)(∀v ∈ V )
a∈δv+ a∈δv− 0, otherwise

dr yar
P P
≤ ca + Cba ub (∀a ∈ A)
r∈R b∈Ba

Bt
P P
kb ub ≤
b∈B t∈T
P
ub + ub0 ≤ 1 (∀b ∈ B)
b0 ∈E b
P
ub ≤ ub0 (∀b ∈ B)
b0 ∈Pb

ub ∈ {0, 1} (∀b ∈ Ba )
yar ∈ [0, 1] (∀r ∈ R)(∀a ∈ A)
zr ∈ [0, 1] (∀r ∈ R).

This optimization problem will act as a subproblem in the solution of the multi-period
expansion models, which are introduced in the next section. It will be used to determine a
suitable preselection of upgrades for further consideration in the two algorithms developed
in Chapter 4.
3.4 Multi-Period Approach 59

3.4. Multi-Period Approach


In this section, we present two models for the multi-period approach which take the for
of multi-period network expansion problems (see Section 2.3.1). In this setting, a feasible
expansion plan must state, which upgrade is nished up to which period, respecting their
implementation times and the available budget in each period of the planning horizon.
Furthermore, a feasible routing of the demand must be found for each period such that
the total prots of the RTCs are maximized.

The above problem description leads to the following optimization problem:

max Sum of (Transport revenues - Transport costs) over all periods

s.t. Route the accepted demand from origin to destination in each period
Respect the track capacities in each period
Keep to the available budget in each period
Respect the construction time of each upgrade
Respect the interdependencies between the upgrades

This multi-period network expansion problem will be referenced as (MNEP) for the rest of
this thesis. In the next two sections, we derive two equivalent models for this problem.

3.4.1. Model (FMNEP)

In this section, we derive a rst model for (MNEP), which will be denoted by (FMNEP)
as it is based on variables which state in which period each upgrade is nished if chosen.
These variables are easy to link to the set of status variables which state if the eect of a
given upgrade is available in a certain period.

The Variables

We dene a set of variables xtb ∈ {0, 1} which answer the question if the construction of
upgrade b ∈ B is nished in period t ∈ T , i.e. if t is the last period of its construction
phase. Accordingly, its eect will be available from that period t on for which

xt−1
b =1 (3.7)

holds.

To be able to decide whether the eect of upgrade b ∈ B is already in place in period t ∈ T̄


(and all following periods), we use auxiliary variables utb ∈ {0, 1}. They are coupled with
the x-variables by introducing the following constraint:

ut+1
b − utb ≤ xtb (∀t ∈ T )(∀b ∈ B). (3.8)
60 3. Modelling the Expansion Problem

It can literally be described as A change in the status of an upgrade is only possible if


the upgrade has been completed up to the previous period. Thus, this constraint ensures
that variable utb can be set to 1 if and only if upgrade b has been completed in any period
preceding t. That way, Condition (3.7) for the availability of the capacity eect is enforced.
As we will see later, a further examination of this coupling constraint will give rise to a
more compact overall model formulation (see Section 4.2).

Furthermore, we have to enforce that the eect of an upgrade stays in place in the subse-
quent periods once it was made available. Therefore, we have to require

utb ≤ ut+1
b (∀t ∈ T̄ )(∀b ∈ B). (3.9)

This constraint keeps the eect of an upgrade from being enabled and reverted arbitrarily.

What we still have to add is the requirement to respect the construction time of each
upgrade. This can easily be achieved by setting

utb = 0 (∀b ∈ B)(∀t ∈ T̄ , t ≤ mb ), (3.10)

such that the coresponding variables actually do not need to be considered. In order to
maintain a more readable representation of the remaining constraints in this exposition,
however, we explicitly include Constraint (3.10) in the model.

Variables yart ∈ [0, 1] determine the routing of the demand in the form of a multi-period
multi-commodity ow. They represent the fraction of relation r ∈ R in period t ∈ T̄ using
track a ∈ A.

Variables z rt ∈ [0, 1] state which fraction of relation r∈R is accepted in period t ∈ T̄ .

The Objective Function

The straightforward adaption of Objective Function (3.1) from Model (NEP) is given by
summing up the individual prots of each time period. Thus, the total prot over the
planning horizon is calaculated as follows:

XX XX X
f1 drt Lr z rt − f2 drt la yart . (3.11)
t∈T r∈R t∈T r∈R a∈A

Now, the objective function also has to consider the observation horizon, which enables a
fair evaluation of the upgrades constructed late in the planning horizon. This is done by
adding the term

!
X X X
r,T +1 r r,T +1 r,T +1
W· f1 d L z − f2 d la yar,T +1
r∈R r∈R a∈A

to (3.11) to obtain the full objective function.


3.4 Multi-Period Approach 61

To simplify the notation, we introduce two time dependent cost factors f1t and f2t , t ∈ T̄ ,
as

fi , if t∈T
fit =
W · fi , if t=T +1

for i ∈ {1, 2}. Thus, we can express the total prot as

X X X X X
f1t drt Lr z rt − f2t drt la yart . (3.12)
t∈T̄ r∈R t∈T̄ r∈R a∈A

The maximization of Function (3.12) is the objective of Model (FMNEP).

The Constraints

The constraints of Model (FMNEP) are the following.

The Routing The routing is modelled as a multi-period ow conservation constraint. In


period t ∈ T̄ , the fraction of relation r∈R emerging from Or is z rt . The same amount
of ow has to arrive at Dr . For the rest of the nodes, the inow is equal to the outow.
Accordingly, this constraint is formulated as

z rt , v = Or

X X  if
yart − yart = −z rt , if v = Dr (∀t ∈ T̄ )(∀r ∈ R)(∀v ∈ V ). (3.13)
0,

a∈δv+ a∈δv− otherwise

The Track Capacities The available track capacities now have to be respected in each
period. Recall that the eect of an upgrade becomes available starting from the period
after its completion. Using variables u for the availability of the upgrades, respecting the
track capacities can be formulated as:

X X
drt yart ≤ ca + Cba utb (∀t ∈ T̄ )(∀a ∈ A). (3.14)
r∈R b∈Ba

The Budget In every period t∈T of the planning horizon, we receive a certain budget
Bt for investment in the network. We consider two alternatives concerning the budget
policy. In the rst setting, unused budget from any period may not be transferred to later
periods (and has to be given back). In the second setting, unused budgets may be used in
subsequent periods.
62 3. Modelling the Expansion Problem

In both cases, we rst need to determine the amount Kbt which is spent on upgrade b∈B
in period t∈T. It is given by:

min(t+mb −1,T )
kb X 0
Kbt = xtb ,
mb
t0 =t

as the construction cost kb is equally distributed over the construction time mb . The sum
in the above expression is equal to one if and only if the upgrade is completed in a period
from the set {t, t + 1, . . . , min(t + mb − 1, T )}, which means that construction has to take
place in period t and the construction cost is incurred.

Thus, we can formulate the budget restriction in the case of a non-transferable budget
as:
X
Kbt ≤ B t (∀t ∈ T ),
b∈B

which, according to the above, is equivalent to:

min(t+mb −1,T )
X kb X 0
xtb ≤ B t (∀t ∈ T ). (3.15)
mb
b∈B t0 =t

In the case of a transferable budget, we demand that the money spent up to each period may
not exceed the total budget received so far. This is expressed by the following constraint,
which is alternative to Constraint (3.15):

0 0
XX X
Kbt ≤ Bt (∀t ∈ T ).
t0 ≤t b∈B t0 ≤t

This can also be written as

min(t0 +mb −1,T )


X X kb X 00
X 0
xtb ≤ Bt (∀t ∈ T ). (3.16)
0
mb
t ≤t b∈B t00 =t0 t0 ≤t

According to the discussion in Section 3.1.3, we stick to the modelling approach of a


transferable budget.

Interdependencies between the Upgrades The case of upgrades whose choice excludes
that of others may be handled similiar to the single-period case. For the subset of upgrades
which are incompatible to upgrade b ∈ B, denoted by Eb ⊂ B , we consider the following
multiple-choice constraint:

X
uTb +1 uTb0+1 ≤ 1 (∀b ∈ B). (3.17)
b0 ∈E b

Due to Monotonicity Constraint (3.9), it suces to consider the observation period T +1.
3.4 Multi-Period Approach 63

For an upgrade b∈B which needs other upgrades as a prerequisite, we have to ensure that
these upgrades, denoted by the set Pb ⊂ B , are nished before starting its construction.
This can be formulated as the following constraint:

utb ≤ ut−m
b0
b
(∀b ∈ B)(∀b0 ∈ Pb )(∀t ∈ T̄ , t > mb ). (3.18)

Complete Statement of (FMNEP)

Model (FMNEP) maximizes the RTCs' prot over all time periods given by Objective
Function (3.12). It incorporates Flow Conservation Constraint (3.13), Capacity Con-
straint (3.14), a transferable budget modelled by Budget Constraint (3.16), Multiple-Choice
Constraint (3.17), Precedence Constraint (3.18), Linking Constraint (3.8), Monotonicity
Constraint (3.9) and Construction Time Constraint (3.10). Its complete statement reads
as follows:

f1t drt Lr z rt − f2t drt la yart


P P P P P
max
t∈T̄ r∈R t∈T̄ r∈R a∈A

z rt , if v = Or

(∀t ∈ T̄ )(∀r ∈ R)

yart yart −z rt , if v = Dr
P P
s.t. − =
a∈δv+ a∈δv−
(∀v ∈ V )
0, otherwise

drt yart Cba utb


P P
≤ ca + (∀t ∈ T̄ )(∀a ∈ A)
r∈R b∈Ba
min(t0 +m
Pb −1,T )
kb 00 0
xtb Bt
P P P
mb ≤ (∀t ∈ T )
t0 ≤t b∈B t00 =t0 t0 ≤t

uTb +1 + uTb0+1 ≤ 1
P
(∀b ∈ B)
b0 ∈Eb
(∀b ∈ B)(∀b0 ∈ Pb )
utb ≤ ut−m
b0
b
(∀t ∈ T̄ , t > mb )
ut+1
b − utb ≤ xtb (∀t ∈ T )(∀b ∈ B)

utb ≤ ut+1
b (∀t ∈ T )(∀b ∈ B)
(∀b ∈ B)
utb = 0
(∀t ∈ T̄ , t ≤ mb )
utb ∈ {0, 1} (∀t ∈ T̄ )(∀b ∈ B)

xtb ∈ {0, 1} (∀t ∈ T )(∀b ∈ B)


(∀r ∈ R)(∀t ∈ T̄ )
yart ∈ [0, 1]
(∀a ∈ A)
z rt ∈ [0, 1] (∀r ∈ R)(∀t ∈ T̄ ).

We will derive a more compact representation of this model in Section 4.2.


64 3. Modelling the Expansion Problem

3.4.2. Model (BMNEP)

In the following, we develop a second model for Problem (MNEP), named (BMNEP) which
is based on a dierent choice of the construction variables. Rather than modelling the end
of the construction phase, they answer the question, if a given upgrade is being built in a
certain period. This modelling alternative might be considered as the more intuitive one,
especially when it comes to stating the budget constraint. However, we will be able to show
in Section 4.1 that the arising model is inferior to Model (FMNEP), which we derived in
the previous section. This is mostly due to its weaker coupling between the construction
variables and the state variables.

The Variables

Our second model (BMNEP) introduces variables gbt ∈ {0, 1} which state whether upgrade
b∈B is currently under construction in period t ∈ T . This way, we know that the eect
of an upgrade b is available from period t on, if

0
X
gbt = mb
t0 ∈T :
t−mb ≤t0 ≤t−1

holds. The availability of upgrade b∈B in period t ∈ T̄ is modelled by the same status
variables utb as in (FMNEP). They are equal to 1 if the corresponding upgrade is available
in that period and 0 otherwise. These status variables have to be linked to the construction
variables g by demanding

0
ut+1
b − utb ≤ gbt (∀b ∈ B)(∀t ∈ T , t ≥ mb )(∀t0 ∈ T , t − mb + 1 ≤ t0 ≤ t), (3.19)

as a change in the status of an upgrade can only occur if its construction was nished in
the period before. An alternative to Constraint (3.19) is its aggregated variant

t
0
X
mb (ut+1
b − utb ) ≤ gbt (∀b ∈ B)(∀t ∈ T̄ , t ≥ mb ).
t0 =t−mb +1

The more compact statement has the disadvantage that it leads to a worse LP bound of
the model. This disadvantage is to be weighted against using a formulation with fewer
constraints. We choose the disaggregated formulation, as the values mb are small in real-
world data sets. Furthermore, this block of constraints is negligible in size against the
routing block of the problem. Thus, we prot from the improved bound at the price of a
small increase of the model size.

The status variables utb have to obey the same monotocity constraint as in (FMNEP) and
cannot be set to 1 for t ≤ mb . Thus, Constraints (3.9) and (3.10) are part of this model,
too.
3.4 Multi-Period Approach 65

The variables yart ∈ [0, 1] for the routing of the demand and z rt ∈ [0, 1] for rejecting portions
of it are the same as in (FMNEP). The former states which fraction of relation r ∈ R uses
track a ∈ A in period t ∈ T̄ , while the latter states which fraction of relation r ∈ R is
accepted in period t ∈ T̄ .

The Objective Function

The objective function of (BMNEP) is the same as in (FMNEP): We maximize the prot
of the RTCs by maximizing Function (3.12).

The Constraints

The following constraints are part of Model (BMNEP):

The Routing and the Track Capacities The routing and the track capacities are mod-
elled by the same variables as in (FMNEP). Therefore, Routing Constraint (3.13) and
Capacity Constraint (3.14) are part of this model, too.

The Budget The main dierence between this model and (FMNEP) lies in the formula-
tion of the budget constraint. As variables gbt state if upgrade b∈B is under construction
in period t ∈ T, it can be stated straightforwardly as

X kb
gt ≤ B t (∀t ∈ T )
mb b
b∈B

in the case of a non-transferable budget, reecting that the costs of each upgrade are
assumed to be evenly distributed over the periods of the construction interval.

For a transferable budget - the case considered here - we take into account the money
received up to each period and subtract what has been spent so far. This yields the
constraint:
X X kb 0 X 0
gt ≤ Bt (∀t ∈ T ). (3.20)
0
mb b 0
t ≤t b∈B t ≤t

Interdependencies between the Upgrades To represent mutual exclusions and prece-


dences among the upgrades, we use Constraints (3.17) and (3.18) from Model (FMNEP).

Complete Statement of (BMNEP)

Altogether, Model (BMNEP) maximizes Objective Function (3.12) under Routing Con-
straint (3.13), Capacity Constraint (3.14), Budget Constraint (3.20), i.e. the transferable
version of the budget, Multiple-Choice Constraint (3.17), Precedence Constraint (3.18),
Linking Constraint (3.19), Monotonicity Constraint (3.9) and Construction Time Con-
66 3. Modelling the Expansion Problem

straint (3.10). The complete model then reads as follows:

f1t drt Lr z rt − f2t drt la yart


P P P P P
max
t∈T̄ r∈R t∈T̄ r∈R a∈A

z rt , if v = Or


yart yart −z rt , if v = Dr
P P
s.t. − = (∀v ∈ V )(∀t ∈ T̄ )(∀r ∈ R)
a∈δv+ a∈δv− 0, otherwise

drt yart Cba utb


P P
≤ ca + (∀t ∈ T̄ )(∀a ∈ A)
r∈R b∈Ba
kb t 0 0
Bt
P P P
m b gb ≤ (∀t ∈ T )
t0 ≤t b∈B t0 ≤t

uTb +1 + uTb0+1 ≤ 1
P
(∀b ∈ B)
b0 ∈E b
(∀b ∈ B)(∀b0 ∈ Pb )
utb ≤ ut−m
b0
b
(∀t ∈ T̄ , t > mb )
0 (∀b ∈ B)(∀t ∈ T , t ≥ mb )
ut+1 − utb ≤ gbt
b (∀t0 ∈ T , t − mb + 1 ≤ t0 ≤ t)
utb ≤ ut+1
b (∀t ∈ T )(∀b ∈ B)
utb = 0 (∀b ∈ B)(∀t ∈ T̄ , t ≤ mb )
gbt ∈ {0, 1} (∀t ∈ T )(∀b ∈ B)
utb ∈ {0, 1} (∀t ∈ T̄ )(∀b ∈ B)
yart ∈ [0, 1] (∀r ∈ R)(∀t ∈ T̄ )(∀a ∈ A)
z rt ∈ [0, 1] (∀r ∈ R)(∀t ∈ T̄ ) .

The following chapter begins with a comparison of Models (FMNEP) and (BMNEP).
4. Model Analysis and Solution
Approaches

In the previous chapter, we were concerned with modelling the optimal expansion of a
railway network. Of particular interest to us is the case of multi-period expansion, i.e. the
task of determining an optimal set of track upgrades together with an optimal schedule
to implement them. In the following, we will treat the development of ecient solution
methods for this problem.

The rst step on this way will be a comparison of the two equivalent models (FMNEP)
and (BMNEP), derived in Sections 3.4.1 and 3.4.2 respectively. We will be able to show
that (FMNEP) provides the better LP bounds and thus is preferrable from a theoretical
point of view. Therefore, all further algorithmic considerations will be based on this model.
They start with the derivation of a more compact representation of Model (FMNEP),
which will be denoted by (CFMNEP). We will see that the structure of its objective func-
tion allows for a strong preprocessing to eliminate routing variables and ow conservation
constraints.

Model (CFMNEP) will then be the basis for an ecient decomposition approach for the
problem, which we call multiple-knapsack decomposition ( MKD for short). In its basic
variant, it is a heuristic algorithm whose underlying idea is to determine a suitable set of
candidate upgrades and to solve a subproblem for the nal choice and scheduling. The
choice of the candidate upgrades is done via solution of the single-period problem (NEP)
for the demand pattern of the observation horizon in order to determine a desirable target
network. This is the master problem of our decomposition approach. The scheduling
subproblem then takes the form of a multiple-knapsack problem that weighs the costs of
the upgrades against their estimated benets, thus the naming of the method.

Finally, we will be able to show that MKD can be extended to an exact method for the
problem via a small modication that allows a suitable embedding into a specialized Ben-
ders decomposition scheme. The result is an iterative variant of MKD that is guaranteed
to converge to the optimal solution. Its fundamental idea can easily be transferred to more
general problem settings in multi-period network design, way beyond the application to
railway networks.

4.1. Comparing Models (FMNEP) and (BMNEP)


The main dierence between our two models (FMNEP) and (BMNEP) lies in the formula-
tion of the budget constraint, or more exactly, in the computation of how much money has

© Springer Fachmedien Wiesbaden 2016


A. Bärmann, Solving Network Design Problems via Decomposition,
Aggregation and Approximation, DOI 10.1007/978-3-658-13913-1_5
68 4. Model Analysis and Solution Approaches

been spent up to a certain period t∈T. In the following, we will prove that (FMNEP) is
better then (BMNEP) from a theoretical point of view in the sense that it provides better
LP bounds.

Theorem 4.1.1. Let LP(FMNEP) and LP(BMNEP) be the LP relaxations of Models


(FMNEP) and (BMNEP) respectively. In addition, let the optimal values of these two LP
relaxations be denoted by OPT(LP(FMNEP)) and OPT(LP(BMNEP)) respectively. Then
there is always
OPT(LP(FMNEP)) ≤ OPT(LP(BMNEP)).

Proof. We prove the above statement by showing that any feasible solution to LP(FMNEP)
can be used to construct a feasible solution to LP(BMNEP) with the same objective value.
Let (x̂, û, ŷ, ẑ) be a feasible solution to LP(FMNEP). Variables u, y and z are part of
both models have the same interpretation in both cases. Thus, to construct an equivalent
solution for LP(BMNEP), we assign them the same values as in the above solution to
LP(FMNEP). From this choice, it is already clear that the new solution will lead to the
same objective value in LP(BMNEP), as only variables y and z contribute to the objective
function, which is equal for both models.

It is left to show that we can nd a suitable choice for variables g to come to a feasible
min(t+mb ,T +1)
solution to LP(BMNEP). Let ĝbt := ûb − ûtb for b ∈ B and t ∈ T . We rst prove
that Budget Constraint (3.20) is fullled. For every t ∈ T , we have:

X X kb 0 X kb X 0 X kb X min(t0 +m ,T +1) 0
ĝ t = ĝbt = ûb b
− ûtb .
0
mb b mb 0 mb 0
t ≤t b∈B b∈B t ≤t b∈B t ≤t

Expanding the inner sum by inserting 0's yields:

0
X kb X min(t0 +m ,T +1) 0
Xb −1 min(t00 +1,T +1)
X kb X t +m 00
ûb b
− ûtb = (ûb − ûtb ),
mb 0 mb 0 00 0
b∈B t ≤t b∈B t ≤t t =t

which can be bounded from above using Linking Constraint (3.8) of (FMNEP) to show
the validity of the budget constraint:

0 0
Xb −1 min(t00 +1,T +1)
X kb X t +m 00
Xb −1 00 X 0
X kb X t +m
(ûb − ûtb ) ≤ x̂tb ≤ Bt .
mb 0 00 0
mb 0 00 0 0
b∈B t ≤t t =t b∈B t ≤t t =t t ≤t

It remains to show that the newly constructed solution (ĝ, û, ŷ, ẑ) fulls Linking Con-
straint (3.19) of (BMNEP) between u and g. Let b ∈ B , t ∈ T̄ with t ≥ mb and
t0 ∈ {t − mb + 1, t − mb + 2, . . . , t}. We nd:

t0 min(t0 +m b ,T +1) 0 0 0
ĝb = ûb − ûtb ≥ ûtb +1 − ûtb ,

where the last inequality is implied by Monotonicity Constraint (3.9). Altogether, this
proves that (ĝ, û, ŷ, ẑ) is a feasible solution to LP(BMNEP) which attains the same objec-
tive value as solution (x̂, û, ŷ, ẑ) to LP(FMNEP).
4.1 Comparing Models (FMNEP) and (BMNEP) 69

The above theorem shows that the LP relaxation of (FMNEP) always provides an upper
bound for the multi-period network design problem which is at least as good as that
of (BMNEP). The following example shows that there are cases where the former provides
a strictly better bound than the latter.

Example 4.1.2. We consider an instance of Problem (MNEP) which is made up as follows.


It consists of two stations, V = {1, 2}, which are connected by one track joining them,
i.e. A = {(1, 2)}. Furthermore, it possesses one relation from node 1 to node 2, i.e.
R = {(1, 2)}, and one upgrade for track (1, 2), i.e. B = {(1, 2)}. The full data for this
instance is given in Table 4.1.
Solution of the LP relaxations of (FMNEP) and (BMNEP) yields the results stated in
Table 4.2. The LP relaxation of (FMNEP) attains an optimal value of 740, while that
of (BMNEP) yields an optimal value of 760 (to compare: The integer optimal value of both
models is 680.). We see that for this example, the LP bound from model (FMNEP) is by
2.7 % better than that of (BMNEP). The reason for this behaviour is the weaker coupling
between the u-variables and the auxiliary g -variables, which leads to a lower estimation of
the expenses for implementing the upgrades in the budget constraint.
Note that Models (FMNEP) and (BMNEP) coincide for instances where all upgrades have
a construction time of one period. Furthermore, it is easy to see from their respective
budget constraints that their LP bound is equal for all instances with a planning horizon
of two periods. Thus, the above example with a planning horizon of three periods (plus the
observation period) is a minimal example where the two bounds dier from each other.

Parameter Value

Planning horizon T =3
Observation horizon W =4
Annual budget Bt = 2 for t∈T
Set of stations V = {1, 2}
Set of tracks A = {(1, 2)}
Track length l(1,2) = 1
Initial track capacity c(1,2) = 2
Set of relations R = {(1, 2)}
Demand d(1,2),t = 10 for t ∈ T̄
Set of upgrades B = {(1, 2)}
Capacity eect C(1,2),(1,2) = 5
Construction cost k(1,2) = 5
Construction time m(1,2) = 2
Revenue per tkm f1 = 50
Cost per tkm f2 = 30

Table 4.1.: Data for the instance of Example 4.1.2


70 4. Model Analysis and Solution Approaches

LP optimal variable values (FMNEP) (BMNEP)

(ḡ 1 , ḡ 2 , ḡ 3 ) - (0.8, 0.8, 0.8)


(x̄1 , x̄2 , x̄3 ) (0.0, 0.6, 0.4) -
(ū1 , ū2 , ū3 , ū4 ) (0.0, 0.0, 0.6, 1.0) (0.0, 0.0, 0.8, 1.0)
(ȳ 1 , ȳ 2 , ȳ 3 , ȳ 4 ) (0.2, 0.2, 0.5, 0.7) (0.2, 0.2, 0.6, 0.7)
(z̄ 1 , z̄ 2 , z̄ 3 , z̄ 4 ) (0.2, 0.2, 0.5, 0.7) (0.2, 0.2, 0.6, 0.7)

LP optimal value 740 760

Integer optimal value 680

Table 4.2.: Results for the instance of Example 4.1.2

The above reasoning leads to the conclusion that (FMNEP) is the better model for Prob-
lem (MNEP) from a theoretical point of view. Its LP bound is always at least as good
as that of (BMNEP) and it is easy to construct examples where the bound provided by
Model (BMNEP) is strictly worse. Therefore, our focus in the remainder of this chapter
is to develop ecient solution strategies based on Formulation (FMNEP).

4.2. A Compact Reformulation of Model (FMNEP)


In this section, we derive a more compact formulation of Model (FMNEP) by eliminating
the x-variables for the nal period of construction of each upgrade. The x-variables are
linked to the state variables u-variables, which reect the availability of the eect of a given
upgrade, by Linking Constraint (3.8). It is is restated here:

ut+1
b − utb ≤ xtb (∀t ∈ T )(∀b ∈ B).

The following consideration shows that the above constraint can be strengthened to an
equality.

Theorem 4.2.1. If Model (FMNEP) possesses an optimal solution, then there exists an
optimal solution which fulls Linking Constraint (3.8) with equality.

Proof. Let (x̄, ū, ȳ, z̄) be an optimal solution to (FMNEP). Monotonicity Constraint (3.9)
ensures that
0 ≤ ūt+1
b − ūtb (∀t ∈ T )(∀b ∈ B)
t
holds. Thus, in case x̄b = 0 holds for some given t ∈ T and b ∈ B , Linking Constraint (3.8)
t t+1
is trivially satised with equality. Now, let x̄b = 1. If ūb = 0, we obtain a feasible
solution to (FMNEP) with an objective value at least as good as that of (x̄, ū, ȳ, z̄) by
switching ut+1
b to 1 and retaining the other variable values. Therefore, we can assume
0
ūt+1
b = 1. With a similar argument, we can assume x̄tb = 0 for 1 ≤ t0 < t . Thus, by
t
Constraints (3.8) and (3.9), we derive ūb = 0. Altogether, this proves the existence of an
optimal solution to (FMNEP) which fulls Linking Constraint (3.8) with equality.
4.2 A Compact Reformulation of Model (FMNEP) 71

Theorem 4.2.1 shows that our formulation (FMNEP) of Problem (MNEP) can be strength-
ened by replacing Linking Constraint (3.8) by

ut+1
b − utb = xtb (∀t ∈ T )(∀b ∈ B).

This allows us to eliminate the x-variables in Budget Constraint (3.16), which is restated
here:
min(t0 +mb −1,T )
X X kb X 00
X 0
xtb ≤ Bt (∀t ∈ T ).
0
mb
t ≤t b∈B t00 =t0 t0 ≤t

For t∈T and xed b ∈ B, its inner summation can be rewritten by evaluating the arising
telescope sum:

min(t0 +mb −1,T ) min(t0 +mb −1,T )


00 00 +1 00 min(t0 +mb ,T +1) 0
X X
xtb = (utb − utb ) = ub − utb .
t00 =t0 t00 =t0

The interpretation of this term is that we have to pay the annual construction cost of
some upgrade b ∈ B if its state of availability changes in the course of the time interval
{t, t + 1, . . . , min(t + mb , T + 1)}. The left-hand side of (3.16) now reads

X X kb min(t0 +m ,T +1) 0
X 0
(ub b
− utb ) ≤ Bt (∀t ∈ T ).
0
m b 0
t ≤t b∈B t ≤t

This can be simplied further by shrinking the new telescope sum. Let t∈T and b∈B
xed. For t ≤ mb , we have

t+m
Xb
min(t0 +mb ,T +1) 0 min(t0 +mb ,T +1) min(t0 ,T +1)
X X
(ub − utb ) = ub = ub
t0 ≤t t0 ≤t t0 =t+1

0
because Construction Time Constraint (3.10) forces utb to be 0 for t0 ≤ mb . For t > mb ,
again using (3.10), we obtain

t
min(t0 +mb ,T +1) 0 min(t0 +mb ,T +1) 0
X X X
(ub − utb ) = ub − utb ,
t0 ≤t t0 ≤t t0 =mb +1

which can be reduced to

t t+m
Xb t t+m
Xb
min(t0 +mb ,T +1) 0 min(t0 ,T +1) 0 min(t0 ,T +1)
X X X
ub − utb = ub − utb = ub .
t0 ≤t t0 =mb +1 t0 =mb +1 t0 =mb +1 t0 =t+1

Thus, we have shown that the result is the same for both t ≤ mb and t > mb and we obtain
a new formulation of the budget constraint in the u-variables:

X kb t+m
Xb min(t0 ,T +1) X 0
ub ≤ Bt (∀t ∈ T ). (4.1)
mb 0 0
b∈B t =t+1 t ≤t
72 4. Model Analysis and Solution Approaches

Literally, the new left-hand side says: In period t, the annual construction cost of any
upgrade b ∈ B has so far been paid as many times as there are periods in {t + 1, t +
2, . . . , t + mb } in which the eect of the upgrade is available. Of course, whenever t + τ
for τ ∈ {1, 2, . . . , mb } refers to a period beyond the scope of T̄ , we have to count the state
in the last period T + 1 several times.

Taking the above derivations together, we obtain a new formulation for Problem (MNEP),
which we denote by (CFMNEP). It maximizes Objective Function (3.12) under Flow
Conservation Constraint (3.13), Capacity Constraint (3.14), Budget Constraint (4.1) for
a transferable budget, Multiple-Choice Constraint (3.17), Precedence Constraint (3.18),
Monotonicity Constraint (3.9) and Construction Time Constraint (3.10). Altogether, it
reads:

f1t drt Lr z rt − f2t drt la yart


P P P P P
max
t∈T̄ r∈R t∈T̄ r∈R a∈A

z rt , if v = Or

(∀t ∈ T̄ )(∀r ∈ R)

yart − yart = −z rt , if v = Dr
P P
s.t.
a∈δv+ a∈δv−
(∀v ∈ V )
0, otherwise

drt yart Cba utb


P P
≤ ca + (∀t ∈ T̄ )(∀a ∈ A)
r∈R b∈Ba
t+m
Pb
kb min(t0 ,T +1) 0
Bt
P P
mb ub ≤ (∀t ∈ T )
b∈B t0 =t+1 t0 ≤t

uTb +1 + uTb0+1 ≤ 1
P
(∀b ∈ B)
b0 ∈E b
(∀b ∈ B)(∀b0 ∈ Pb )
utb ≤ ut−m
b0
b
(∀t ∈ T̄ , t > mb )
utb ≤ ut+1
b (∀t ∈ T )(∀b ∈ B)
(∀b ∈ B)
utb = 0
(∀t ∈ T̄ , t ≤ mb )
utb ∈ {0, 1} (∀t ∈ T̄ )(∀b ∈ B)
(∀r ∈ R)(∀t ∈ T̄ )
yart ∈ [0, 1]
(∀a ∈ A)
z rt ∈ [0, 1] (∀r ∈ R)(∀t ∈ T̄ ).

We will use this more compact formulation in the following to derive solution algorithms
for Problem (MNEP).

4.3. Preprocessing of (NEP) and (CFMNEP)


The number of variables in Models (NEP) and (CFMNEP) is vastly dominated by the
routing decisions, i.e. the y -variables. In a realistic data setting, they account for about
99 % of the variables. It is thus desirable to eliminate as many of them as possible in
preprocessing. The structure of the objective functions of Models (NEP) and (CFMNEP)
4.3 Preprocessing of (NEP) and (CFMNEP) 73

allows for a very ecient preprocessing. In the following, we present this method for (NEP).
It applies to (CFMNEP) in a similar fashion.

Objective Function (3.1), which is maximized in (NEP), reads

X X X
f1 dr Lr z r − f2 dr la yar .
r∈R r∈R a∈A

It evaluates the prot of the RTCs from the transportation of the demand. For each single
relation, it is possible to state whether a given transportation path is protable or not.
The revenues for a relation r ∈ R are proportional to the length of a shortest path in the
network connecting nodes Or and Dr by a factor of f1 . The expenses on the other hand
are proportional to the length of the path which the relation actually takes by a factor of
f2 . Therefore, it is uneconomic to transport any relation along a path for which the ratio
between its length and that of a shortest path is higher than the ration between f1 and
f2 . Instead of transporting the relation along such a detour, this transport order would be
rejected in an optimal solution to (NEP).

We can use this property to exclude a high number of uneconomic routings in preprocessing.
As (NEP) is an arc ow model, we need a criterion to determine which arcs a given relation
will not use in an optimal solution. A criterion which is easy to check in advance is the
following:

Proposition 4.3.1. Let r ∈ R and a = (v, w) ∈ A. Furthermore, let Λrv be the length of
a shortest path from Or to v , and let Λrw be the length of a shortest path from w to Dr . If

f1 r
Λrv + la + Λrw ≥ L
f2

holds, we can set variable yar in (NEP) to 0 and eliminate it from the problem.

The eciency of this straightforward rule depends on the actual ratio between f1 and f2 .
The closer together the two are, the more variables can be eliminated beforehand.

A similar rule allows for the elimation of unnecessary routing constraints:

Proposition 4.3.2. Let r ∈ R and v ∈ V . Furthermore, let Λrv,1 be the length of a shortest
path from Or to v , and let Λrv,2 be the length of shortest path from v to Dr . If

f1 r
Λrv,1 + Λrv,2 ≥ L
f2

holds, then Routing Constraint (3.2) corresponding to relation r and node v is redundant
and can be eliminated from (NEP).

The computational results in Section 5.2.2 show that these two preprocessing rules allow
for a signicant reduction in problem size for real-world data settings.
74 4. Model Analysis and Solution Approaches

4.4. Decomposition Algorithms

Real-world problem instances for the expansion of railway networks are usually too large to
solve with standard branch-and-bound solvers, despite the model enhancements developed
so far. The ambition persued in this thesis is the solution of a Germany-wide instance over
a planning horizon of 20 years. The resulting (preprocessed) model size of (CFMNEP) is
not only a problem for the algorithms of a standard solver, it already poses problems to
store the model in the memory of an up-to-date workstation. Therefore, we develop two
solution approaches to deal with this optimization problem in the following, a heuristic
decomposition scheme and its exact extension.

4.4.1. Multiple-Knapsack Decomposition

This section presents an approach to divide (CFMNEP) into subproblems whose individual
solutions form a solution to the original problem when put together. The key for an ecient
solution algorithm is to decompose the two aspects of choosing the adequate upgrades for
the expansion of the network and routing the demand. This will allow us to deal with the
routing in each time period of the planning horizon seperately.

The approach developed here consists of the following four steps:

1. Solve the single-period network expansion problem (NEP) to determine a set of


desirable upgrades.

2. Evaluate these candidate upgrades by estimating their protability in each period.

3. Solve a subproblem to determine the optimal choice and schedule of the upgrades
based on the evaluation from Step 2.

4. Determine the optimal routing under the expansion schedule chosen.

The method starts by determining a choice of candidate upgrades, which is done by solv-
ing (NEP). The advantage is a (comparatively) quick-to-obtain preselection among the
available upgrades for the expansion of the network. The motivation to do so is that (NEP)
is based on the demand of the observation horizon, which can be seen as the target demand
pattern to be dealt with. This is all the more true under the consideration of a reasonably
long observation horizon, which results in a higher weight of the corresponding time pe-
riod T + 1. Finally, in a realistic instance, the development of the demand over time can
be assumed to represent rather a converging development over the planning horizon than
a completely chaotic one. This justies the hope to nd an ecient choice of candidate
upgrades from (NEP) with respect to optimization over the whole planning horizon. The
subsequent steps of the algorithm are presented in the following.
4.4 Decomposition Algorithms 75

Evaluating the Upgrades

After the rst step of the decomposition algorithm, we have determined a preselection
of desirable upgrades to expand the network from Problem (NEP). The next step is to
evaluate them according to their protability. For each upgrade and for each planning
period, we would like to estimate the additional prot which could be achieved given the
availability of the resulting additional capacities. The estimated benets will serve as
weights in the scheduling subproblem in the third step of the algorithm.

The way to obtain these estimates is via a period-wise evaluation of the prots which are
achievable in the two cases with and without the implementation of the selected upgrades.
First, we determine the prot of the RTCs in the case that no network expansions are
made. Therefore, we solve the following optimization problems denoted by (FBEt ) for
t ∈ T̄ , to obtain the optimal ow before expansion :

max f1t drt Lr z rt − f2t drt la yart


P P P
r∈R r∈R a∈A

z rt , v = Or

 if
yart yart −z rt , v = Dr
P P
s.t. − = if (∀v ∈ V )(∀r ∈ R)
a∈δv+ a∈δv− 0,

otherwise

drt yart ≤ ca
P
(∀a ∈ A)
r∈R

yart ∈ [0, 1] (∀r ∈ R)(∀a ∈ A)


z rt ∈ [0, 1] (∀r ∈ R).

This problem can be seen as an adaption of (NEP), where t is assumed to be the target
period for the network expansion and where all upgrade decisions are xed to 0. Thus, it
remains the maximization of the prot by choosing the fraction of the demand to accept
for each relation as well as its optimal routing. Let η t , t ∈ T̄ , be the optimal value of
(FBEt ). These values are the prots achievable with the initial network conguration for
each period.

Next, we will put these values in comparison to the prots achievable under the preselection
of upgrades from (NEP). We therefore assume that all candidate upgrades are implemented
from the rst period on and determine their respective degree of utilization. Let B̄a ⊆ Ba
denote the subset of candidate upgrades for track a∈A which was chosen in (NEP). The
term
X
Cba
b∈B̄a

then states the amount of extra capacity made available on track a if all upgrades in B̄a
are realized. Now we determine optimal routings in each period, assuming that these extra
capacities are available from the start. More exactly, we solve the following optimization
76 4. Model Analysis and Solution Approaches

problem denoted by (FAEt ) for t ∈ T̄ to obtain the optimal ow after expansion :

max f1t drt Lr z rt − f2t drt la yart


P P P
r∈R r∈R a∈A

z rt , if v = Or


yart yart −z rt , if v = Dr
P P
s.t. − = (∀v ∈ V )(∀r ∈ R)
a∈δv+ a∈δv− 0, otherwise

drt yart
P P
≤ ca + Cba (∀a ∈ A)
r∈R b∈B̄a

yart ∈ [0, 1] (∀r ∈ R)(∀a ∈ A)


z rt ∈ [0, 1] (∀r ∈ R).

This problem determines an optimal ow in the network given the possibility to use the
candidate upgrades from (NEP). We are expecially interested in the utilization of these
upgrades in an optimal solution (ȳ t , z̄ t ) to (FAEt ), i.e. we would like to know which fraction
of the extra capacity is used on each track in any of the periods. This information is given
by the utilization value sta ∈ [0, 1], which is calculated as follows:

P rt rt !
d ȳa −ca

r∈RP
 max 0, , B̄a 6= ∅

if

Cba
sta := b∈B̄a



0, otherwise.

Basicly, we determine the amount by which the base capacity of each track is exceeded
and divide this value by the available extra capacity, putting a value of zero if no upgrades
were chosen for a track or if the extra capacity is not used. These values will be used to
estimate the additional prot possible by upgrading each track a ∈ A.
t
Let θ for t ∈ T̄ t
be the optimal value of (FAE ). It states the prot which is achievable
in each period using the candidate upgrades. We denote the prot growth in each period
t ∈ T̄ by
∆t := θt − η t .
In the following, the usage of extra capacities in the solution to (FAEt ) will be used to
determine which upgrades are most benecial to realize. More exactly, it is used to estimate
the amount of extra prot which each upgrade contributes to the objective function when
implemented. This is done in two steps. First, we distribute the prot growth in each
period among the tracks, weighted by their utilization values. Let

X
S t := sta
a∈A

be the sum of the utilization values over all tracks in period t ∈ T̄ as obtained from the
solution to (FAEt ). We then dene the estimated prot growth per track as

sta
λta := · ∆t , t ∈ T̄ , a ∈ A.
St
4.4 Decomposition Algorithms 77

The second step is to estimate which extra prot can achieved by the implementation of
each single upgrade on a given track in a given period. This estimated prot growth per
upgrade is dened as

C
P ba
X
µtb := · λt , t ∈ T̄ , b ∈ B̄a .
C b0 a a
a∈Ab
b0 ∈B̄a

Note that each upgrade can have an eect on multiple tracks. The estimates thus sum
over those tracks in a ∈ Ab which are aected by an upgrade b ∈ B̄a . These values µtb are
used as input in the next step of the algorithm.

Final Choice and Scheduling of the Upgrades

The third step of the algorithm makes the nal choice which upgrades from the set of can-
didate upgrades to implement and when. To do so, it weighs the construction costs against
their estimated benets in each period. Furthermore, it has to respect the constraints for
a feasible implementation of the methods. These requirements lead to an optimization
subproblem which can be seen as an adaption of (CFMNEP).

The objective is an optimal choice of upgrades maximizing the estimated extra prot over
planning horizon and observation horizon. Thus, the objective function is given by:

XX
µtb utb .
t∈T̄ b∈B̄a

The implementation of the upgrades is restricted by Budget Constraint (4.1), Multiple-


Choice Constraint (3.17), Precedence Constraint (3.18), Monotonicity Constraint (3.9)
and Construction Time Constraint (3.10) of (CFMNEP). The complete problem reads as
follows:

µtb utb
P P
max
t∈T̄ b∈B̄a
t+m
Pb
kb min(t0 ,T +1) 0
Bt
P P
s.t. mb ub ≤ (∀t ∈ T )
b∈B t0 =t+1 t0 ≤t

uTb +1 + uTb0+1 ≤ 1
P
(∀b ∈ B)
b0 ∈Eb
utb ≤ ut−m
b0
b
(∀b ∈ B)(∀b0 ∈ Pb )(∀t ∈ T , t > mb )
utb ≤ ut+1
b (∀t ∈ T )(∀b ∈ B̄)
utb = 0 (∀b ∈ B)(∀t ∈ T̄ , t ≤ mb )
utb ∈ {0, 1} (∀t ∈ T̄ )(∀b ∈ B̄).

This problem can be stated as a multiple-knapsack problem with additional constraints


by variable substitution. We denote it by (BK) for budget knapsack. Its solution yields a
feasible schedule for the expansion of the network.
78 4. Model Analysis and Solution Approaches

Determining the Solution for (CFMNEP)

The expansion schedule obtained from (BK) can now be extended to a complete solution
to (CFMNEP). Let ū be an optimal solution to (BK). As ū fulls all scheduling constraints
from (CFMNEP), we can x the corresponding expansion schedule in (CFMNEP) and solve
the problem for the remaining variables. In other words, we determine the optimal routing
of the demand given the expansion plan from (BK) to obtain a feasible solution (ū, ȳ, z̄) to
the original problem. That the resulting solution is indeed feasible stems from the fact that
any demand relations can be rejected if the required capacities for their transportation are
not available.

Note that from the view of the planner, this algorithm has a variety of advantages. First,
it is important to note that the solution of each of its subproblems itself yields relevant
information for the planning process. Step 1 determines a selection of upgrades which are
benecial from the view point of the target demand pattern the network shall be adjusted
to. Step 2 estimates for each upgrade to which increase in prot its implementation will
lead. In Step 3, the planner can determine the optimal nal choice of upgrades and their
optimal temporal distribution with respect to the budget as well as technical requirements.
Finally, he is presented the optimal routing of the trains through the network for each time
period in Step 4, which enables him to judge the obtained expansion plan in the light of
the arising trac ows and thus to check its plausibility.

The second advantage of this decomposition approach is that it provides the planner with
the opportunity to inuence its outcome in an interactive fashion. If desired, he can alter
the selection of candidate upgrades after Step 1 to better t his needs. He can also integrate
his own ideas of how protable he considers each upgrade to be to modify the estimations of
Step 2. And he can inuence the nal choice of upgrades and their order of implementation
and then obtains an evaluation of this schedule in the form of an optimal routing which
can compared to that of other upgrade plans.

Last but not least, he obtains the solution in comparatively short time (see Section 5.3),
which is an absolute requirement to be able to compare dierent demand scenarios and
the expansion plans they lead to.

4.4.2. Extending the Heuristic to an Exact Decomposition Method

As we will see in the next chapter, the heuristic developed here for the network expansion
problem for the German railway network yields solutions of very high quality. Thus, from
the point of view of the application, it is almost as good as an exact method  or even better,
taking into account its short solution time. Furthermore, the great amount of inherent
data uncertainty in the demand forecast makes the eort to obtain an optimal solution
pointless from the start. Incorporation of techniques from robust optimization or stochastic
optimization might be a remedy here. On the other hand, planners at GSV stated that
they were less interested in a solution taking into account the eects of uncertainty, but
rather in studying how the solution changes when dierent scenarios are used as input.
4.4 Decomposition Algorithms 79

Altogether, we see no need to actually implement an exact method to the problem. From
the mathematical point of view, however, it is interesting to consider how the heuristic can
be extended to an exact method. Indeed, the multiple-knapsack decomposition developed
in the previous section is only one step away from an exact algorithm. In the following,
we show how this gap can be bridged by relating the method to a specialized Benders
decomposition scheme, which is the framework for such an exact extension. Altogether,
we show that an iterative adaption of the heuristic can be brought to converge to an
optimal solution to the problem.

Optimal Choice of the Upgrades

If we want to obtain an iterative version of MKD which converges to an optimal solution,


two properties have to be satised. Firstly, the choice of the candidate upgrades has to
converge to an optimal choice of upgrades. Secondly, the implementation schedule for the
optimal choice of upgrades has to converge to an optimal schedule. Both properties can
be ensured by embedding the heuristic into a suitable Benders decomposition, which is
described in the following.

Recall that the rst step of the heuristic is to choose a suitable subset of candidate upgrades
to implement. This is done by solving the single-period expansion problem (NEP) with
respect to the demand of the observation horizon as well as the total budget obtained over
the planning horizon. The main idea for an exact extension of the algorithm is that the very
same set of candidate measures would be obtained in the rst iteration of a special Benders
decomposition of (CFMNEP)  namely when projecting out all ow-related variables y and
z for all time periods except for the period representing the observation horizon. Secondly,
the primal Benders subproblems to evaluate the so-computed network expansion are very
similar to those solved by the heuristic. In the rst iteration of the Benders scheme,
they coincide. Via the introduction of Benders optimality cuts, both the choice and the
scheduling of the upgrades can now be improved in each of the subsequent iterations until
optimality can be proved. These ideas are detailed next, starting with the optimal choice
of the upgrades.

An optimal solution to Model (CFMNEP) delivers an optimal upgrade schedule for the
railway network. In particular, the optimal values for variables uTb +1 for b∈B describe an
optimal choice of upgrades. This property of course remains valid when projecting out (part
of ) the continuous variables, as it can be done algorithmically via Benders decomposition.
If we project out all ow variables yt and z t for all periods t ∈ T of the planning horizon,
but not those for the observation period t = T + 1, we obtain an MIP that is very similar
in structure to Model (NEP), which acted as our master problem in the heuristic version
of multiple-knapsack decomposition. It takes the form of a single-period network design
problem for period t=T +1 with further side-constraints to incorporate the information
of earlier time periods as well as the corresponding u-variables. This specialized Benders
decomposition can be initiated with the following restricted master problem, which contains
80 4. Model Analysis and Solution Approaches

none of the additional cutting planes describing the projection:

max f1T +1 dr,T +1 Lr z r,T +1 − f2T +1 dr,T +1 la yar,T +1 + φt


P P P P
r∈R r∈R a∈A t∈T

z r,T +1 , if v = Or


yar,T +1 yar,T +1 −z r,T +1 , if v = Dr
P P
s.t. − = (∀v ∈ V )(∀r ∈ R)
a∈δv+ a∈δv− 0, otherwise

dr,T +1 yar,T +1 Cba uTb +1


P P
≤ ca + (∀a ∈ A)
r∈R b∈Ba
t+m
Pb
kb min(t0 ,T +1) 0
Bt
P P
mb ub ≤ (∀t ∈ T )
b∈B t0 =t+1 t0 ≤t

uTb +1 + uTb0+1 ≤ 1
P
(∀b ∈ B)
b0 ∈Eb
(∀b ∈ B)(∀b0 ∈ Pb )
utb ≤ ut−m
b0
b
(∀t ∈ T̄ , t > mb )
utb ≤ ut+1
b (∀t ∈ T )(∀b ∈ B)
(∀t ∈ T̄ , t ≤ mb )
utb = 0
(∀b ∈ B)
utb ∈ {0, 1} (∀t ∈ T̄ )(∀b ∈ Ba )

yar,T +1 ∈ [0, 1] (∀r ∈ R)(∀a ∈ A)

z r,T +1 ∈ [0, 1] (∀r ∈ R)

φt ∈ [0, M ] (∀t ∈ T ).
(4.2)

Variables φt are the prots of the years t∈T of the planning horizon, whose estimation
depending on the upgrades is iteratively improved via cutting planes. Note that we need
to initialize them with a valid upper bound M. A suitable value may be found via |T |
multi-commodity ow calculations by assuming that all available upgrades in the set B
were present from the start.

This restricted master problem optimizes the choice of upgrades for an optimal routing
in the observation period. Indeed, it is nothing else but an extended formulation for
the single-period problem (NEP). In the rst iteration of Benders decomposition, all the
variables utb for t∈T and b∈B can be set to zero, as the prot estimations φt do not yet
T +1
deliver guidance on how to schedule the upgrades. The values of ub for b ∈ B coincide
with those calculated by (NEP). Let B̄a ⊂ Ba again denote the chosen subset of upgrades
for track a ∈ A.
Altogether, we can replace the solution of (NEP) in the multiple-knapsack decomposition
by the solution of Problem (4.2). In the rst iteration, it will make no dierence. In the
subsequent iterations, the choice of candidate measures will be improved by the more and
more exact estimations for the prots φt . In the end, the choice of the upgrades in this
4.4 Decomposition Algorithms 81

extended multiple-knapsack decomposition will converge to an optimal choice of upgrades


because the Benders master problem will converge to an optimal solution.

Optimal Scheduling of the Upgrades

We just saw that our iterative version of the multiple-knapsack decomposition will be
able to nd an optimal choice of upgrades if this choice is driven by the Benders master
problem (4.2). In the second step, we need to ensure that the method terminates with an
optimal scheduling of the upgrades. To do so, it is necessary to modify Subproblems (FBEt )
and (FAEt ) which evaluate a given choice of upgrades for each period t ∈ T̄ . Furthermore,
we have to adapt Problem (BK) which determines the scheduling of the upgrades.

We begin with Problem (BK). The idea to make the solution to (BK) converge to an
optimal schedule again works by tying it to the solution to the Benders master problem.
We impose that no upgrade may be implemented later in (BK) than in the Benders master.
This is done by adding the constraint

utb ≥ ūtb (∀t ∈ T̄ )(∀b ∈ B).

Especially, this constraint implies that all upgrades chosen in the Benders master are
actually implemented. In other words, they are no longer candidate upgrades. By this step,
we have already ensured convergence to an optimal schedule. However, to accelerate this
convergence and to attain good solutions already in early iterations, we have to make use of
the remaining degree of freedom in Problem (BK), which is to prepone the implementation
of certain upgrades. This requires us to modify the calculation of the estimated benets in
each period t∈T to improve their precision over the subsequent iterations of the method.
For period T + 1, they need not be calculated any more.

Both Problems (FBEt ) and (FAEt ) are closely related to the Benders subproblem for period
t∈T according to the above decomposition, which reads:

max f1t drt Lr z rt − f2t drt la yart


P P P
r∈R r∈R a∈A

z rt , if v = Or


yart − yart = −z rt , if v = Dr
P P
s.t. (∀v ∈ V )(∀r ∈ R)
a∈δv+ a∈δv− 0, otherwise

(4.3)
drt yart Cba ūtb
P P
≤ ca + (∀a ∈ A)
r∈R b∈Ba

yart ∈ [0, 1] (∀r ∈ R)(∀a ∈ A)


z rt ∈ [0, 1] (∀r ∈ R),

where ū is the network expansion plan determined by the master problem. Problem (FBEt )
t
is retained from Problem (4.3) by choosing ū = 0, while Problem (FAE ) is obtained by
t
choosing ūb = 1 for all upgrades b chosen in the master problem for all periods t ∈ T .
82 4. Model Analysis and Solution Approaches

Dierent from the approach developed in Section 4.4.1, we now assume that each chosen
upgrade has to be implemented until the period determined by the master problem at the
latest. The remaining degree of freedom is preponing some of these upgrades. To estimate
the additional prot by doing so, we have to modify the calculation of the parameters µ
in Problem (BK).

Obviously, determining the prot achievable by nishing each upgrade in the same period
as suggested by the master problem leads to Subproblem (4.3). Thus, η t is now given as the
the optimal value of (4.3) for period t∈T instead of determining it via Problem (FBE ).
t

Its new interpretation is the prot obtained in each period without preponing any of the
upgrades.

To estimate the eect of preponing upgrades, we can leave Problem (FAEt ) intact, but
have to adapt the calculation of the utilization values st . For some expansion plan ū from
the master problem and an optimal solution (ȳ t , z̄ t ) to Problem (FAEt ), they are now
dened as

  !
P rt rt
Cba ūtb
P
d ȳa − ca +



 max 0, r∈R
 b∈B̄a
{b ∈ B̄a | ūtb = 0} =
6 ∅
, if

sta := Cba (1−ūtb )
 P

 b∈B̄a



0, otherwise.

That means, we add the extra capacity by those upgrades which have to be nished up
to period t (according to the schedule of the master problem) to the base capacity of each
track. The capacity used in excess of this value is divided by the extra capacity from the
upgrades which can be preponed from later periods. All in all, we obtain the fraction of
preponed extra capacity used in each period.

Using θt , the optimal value of Problem (FAEt ), we can determine ∆t as

∆t := θt − η t .

It is now to be interpreted as the extra prot achievable by preponing upgrades. The


remaining calculations to determine the µ-values are the same as in Section 4.4.1. Thus,
we can now use Problem (BK) to determine a feasible upgrade schedule which can be
expected to be much better than that given by the solution to the master problem in the
early iterations of the algorithm.

Preparing the next Iteration

To initialize the next iteration, we have to update the Benders master problem (4.2). This
is done by adding an optimality cut for each period t ∈ T which is derived from the
objective function of the respective dual subproblem. For a master expansion ū, the dual
4.4 Decomposition Algorithms 83

subproblem for period t∈T is given by

rt − αrt ) + Cba ūtb )βat + γart + ζ rt


P P P P P P
min (αO r Dr (ca +
r∈R a∈A b∈Ba r∈R a∈A r∈R

s.t. αvrt − αw
rt + drt β t + γ rt ≥ −f t drt l
a a 2 a (∀r ∈ R)(∀a = (v, w) ∈ A)
rt
αO − rt
αD + ζ rt ≥ −f1t drt Lr (∀r ∈ R)
r r

βat ≥ 0 (∀a ∈ A)
γart ≥ 0 (∀r ∈ R)(∀a ∈ A)
ζ rt ≥ 0 (∀r ∈ R),

where αt are the dual variables to the ow conservation constraints, β t those to the capacity
constraints and γ t and ζ t those to the upper bounds of y t and z t respectively.
For each t∈T, the optimality cut is then given by

rt − ᾱrt + γ̄art + ζ̄ rt ) + ca β̄at + (Cba β̄at )utb ≥ φt ,


P P P P
(ᾱO r Dr
r∈R a∈A a∈A b∈Ba

for a dual optimal solution (ᾱt , β̄ t , γ̄ t , ζ̄ t ). Adding this cut to Master Problem (4.2) and
resolving it yields a new upgrade schedule to be employed in the subsequent iteration. The
algorithm terminates with an optimal solution to Problem (CFMNEP) when the upper
bound given by the master problem coincides with the value of the solution determined by
the modied multiple-knapsack decomposition.

Summary of the Method and Discussion

In the previous derivation, we saw that a modied variant of MKD yields an exact al-
gorithm for Problem (CFMNEP) by an embedding into a Benders scheme. The resulting
algorithm can be summarized as follows:

1. Solve the Benders master problem (4.2) to obtain an upgrade schedule ū.
2. Solve the Benders subproblem (4.3) for each t∈T to obtain the values η t , the prots
without preponing upgrades.

3. Solve problem (FAEt ) for each t∈T to obtain the values θt , the prots possible by
preponing upgrades, as well as the utilization values st .
4. Use ηt , θt and st to calculate the µ-values and solve Problem (BK).

5. From the upgrade schedule determined by Problem (BK) compute a complete solu-
tion to Problem (CFMNEP).

6. If the objective value of the solution determined in Step 5 coincides with the bound
given by Master Problem (4.2) STOP with an optimal solution. Otherwise use the
dual values of Subproblem (4.3) to compute an optimality cut for period t∈T and
add it to Master Problem (4.2). Go to Step 1.
84 4. Model Analysis and Solution Approaches

Basicly, the modied multiple-knapsack decomposition is executed with a tentative upgrade


schedule from the master problem which is updated in each iteration. Especially, the
rst iteration of the above algorithm is nothing else but the original multiple-knapsack
decomposition when dropping the requirement that all upgrades chosen in the master
problem have to be implemented.

Obvioulsy, this exact algorithm can also be used as an iterative improvement heuristic, as
we may stop at any point with an estimation of the remaining optimality gap. Moreover,
from the view of the planner, the interactive component of the decomposition method is
maintained, as he may still inuence the chosen upgrades and their tentative schedule as
well as the estimation of their protability in Problem (BK).

We did not implement this exact variant of our method, as the original multiple-knapsack
decomposition already yields solutions of very high quality. We will see in the next chapter
that for our reference problems, the highest optimality gap is still below 2 %, which is more
than satisfying from the view of the application. This can be explained by the fact that
the structure of the input data of this real-world problem ts the central assumption of
the method very well, namely that upgrade candidates with higher utilization contribute
more to the increase in prot. There may be situations where this is not the case, e.g.
when considering candidates for newly-built links with largely oversized capacities. Then
the exact algorithm developed in this section may be employed and can be expected to
yield high-quality solutions within few iterations.
5. Case Study for the German Railway
Network

In the previous chapters, we explained the necessary background on railway infrastructure


development, both from the technical and the mathematical point of view, we elaborated
on ecient model formulations for the problem and we developed practical algorithms to
solve the problem. Now, we put together all this knowledge to solve real-world problem
instances from DB Mobility Logistics AG. We begin by describing the software planning
tool we implemented for practical use by our industry partner. In a rst step, it is tested
on subnetworks of the German railway network to evaluate the computational behaviour
of our methods. Then we describe our solution for the whole German railway network,
which is the actually interesting test case. We show that our solution is of very high quality
compared to an optimal solution to the problem and that our solution is achievable with
much less computational eort. We also give a detailed discussion of the network expansion
it proposes and compare it to planning results by experienced planners at GSV.

5.1. Our Planning Software and the Computational Setup

We implemented the single-period planning model (NEP) for the expansion of the Ger-
man railway network as well as the three equivalent multi-period models (FMNEP), its
more compact version (CFMNEP) and the alternative formulation (BMNEP). Further-
more, we implemented the multiple-knapsack decomposition scheme developed in Sec-
tion 4.4.1. These components have been integrated into a software package written in
C++ that we provided to our industry partner DB Mobility Logistics AG. It serves as
a prototype implementation from which their department for trac simulation and prog-
nosis, GSV, can now build a specialized planning software for network expansion studies.
Altogether, it allows planners in the eld to quickly evaluate dierent strategic choices
such as budget allocation and considered upgrades under dierent demand scenarios.

In the following, we demonstrate the exibility and versatility of our software tool and
assess the computational properties of the employed models as well as the eency and
solution quality of our algorithms. To this end, we perform a broad set of computational
experiments for their evaluation under a setup described henceforth. They show the suit-
edness of our approach to solve this real-world problem as well as the benets of addressing
it with the techniques of discrete optimization.

© Springer Fachmedien Wiesbaden 2016


A. Bärmann, Solving Network Design Problems via Decomposition,
Aggregation and Approximation, DOI 10.1007/978-3-658-13913-1_6
86 5. Case Study for the German Railway Network

5.1.1. Implementation
Our software package provides the option to solve the network expansion problem in both
the single- and the multi-period case. This includes an input-output interface to read in the
problem data given in the data format used at GSV and to deliver the solution in a format
in which it can directly be processed in dedicated trac simulation software. Furthermore,
it oers dierent parameter settings for the solution process including the computational
resources to be used. For the multi-period case, it allows for the choice between an exact
solution of the problem via all three models (FMNEP), (CFMNEP) and (BMNEP) as well
as using the heuristic variant of multiple-knapsack decompositon.

We use Gurobi's branch-and-bound implementation (Gurobi Optimization, Inc. (2014)) to


solve the arising linear and mixed-integer linear optimization problems. The models are
built and solved via its C++-Interface. The LEMON Graph Library (Dezs® et al. (2011))
is used to represent the network structures and to perform auxiliary graph computations
such as the shortest path calculations to determine the revenues of the demand orders.

The computations presented here were performed on a compute server with a Six-Core
AMD Opteron— 2435 processor using all 6 cores and 64 GB of memory. The employed
version of Gurobi was 5.5, the lemon version was 1.2.3. All specic parameter settings are
described in the respective sections.

5.1.2. The Input Data


Our industry partner provided us with a high-quality real-world data set which contains
all necessary input data to our models. It includes a macro-scale representation of the
complete German rail freight network. Its set of nodes are all German rail freight stations
at which routing decisions are possible. The arcs are the tracks joining them, where
segments with intermediate stations are contracted if the track properties do not change
signicantly.

The demand data results from a sample day of rail freight trac, 17th December 2010,
as well as a carefully undertaken GSV-internal demand forecast for a representative sam-
ple day in 2030. Thus, we take 2010 as the base year in all computations and consider
a planning horizon T of 20 years. A detailed discussion of this forecast is given in Sec-
tion 5.3.2. To obtain an estimation for the trac in all intermediate years, we use a linear
interpolation between the two boundary values. Note that we assume passenger trac to
remain constant between 2010 and 2030. The observation horizon W, which provides for
the possibility of the amortization of investments undertaken late in the planning horizon,
is also set to be 20 years.

The measures considered here to increase the capacity of the links in the network were
the construction of new tracks, electrication, block-size reduction as well as increasing
the maximal velocity on a link. For a detailed description of these measures, we refer to
Section 2.2.3. Their concretions, namely specic link upgrades, were obtained by employing
a rule-based approach to determine which measures are applicable to which links and by
how much link capacity can be improved by applying them. Both depends heavily on the
5.1 Our Planning Software and the Computational Setup 87

properties of a link such as the line standard, the currently allowable traction, its current
number of tracks and its current block size. These criteria and the necessary capacity
estimations were compiled as well by GSV.

Finally, GSV provided us with suitable choices for the monetary parameters of the problem.
These entail an annual budget for investment Bt
of 700 million Euros for each year t ∈ T ,
a revenue rate f1 of 50 Euros per train kilometre along the shortest path and a transport
cost rate f2 of 30 Euros per train kilometre along the path actually taken by a demand
relation.

Together, these data are the foundation for our Germany-wide case study on an optimal ex-
pansion of the rail freight network until 2030, whose results are presented in Section 5.3.
As for its mathematical properties, this data set leads to a very high-dimensional network
design problem. The underlying directed graph of the network contains about 1600 stations
and 5200 tracks. For each pair of opposite tracks, 2 to 3 available upgrades are to be
considered on average while the number of demand relations is about 3600. These numbers
have to be multiplied by 20 to account for the temporal expansion of the network.

Considering the above gures, it becomes obvious that this full problem instance is not
suitable to systematically assess the computational properties of the models and algorithms
developed here. For this reason, we constructed subinstances of the Germany-wide net-
work which correspond to subnetworks of a single or of multiple adjacent German federal
states. This subdivision is shown in Table 5.1, which shows the federal states belonging to
each subnetwork. Together, they form a complete covering of the whole German railway
network.

Network Involved federal states

SH-HA Schleswig-Holstein, Hamburg


B-BB-MV Berlin, Brandenburg, Mecklenburg-Vorpommern
HE-RP-SL Hessen, Rheinland-Pfalz, Saarland
NRW Nordrhein-Westfalen
BA-BW Bayern, Baden-Württemberg
NS-BR Niedersachsen, Bremen
TH-S-SA Thüringen, Sachsen, Sachsen-Anhalt
Deutschland The entire German instance

Table 5.1.: Layout of the railway networks under consideration

The subnetworks have been dened such that they allow to incorporate the complete
demand scenario of the entire German instance. This was achieved by shrinking all federal
states which are not part of a specic instance to one node only, aggregating also all demand
originating and terminating there to that one node. This way, we obtained subinstances
with the same trac density as in the original instance, preserving most part of the cost
structure, as path lengths are vastly maintained. Also preserved is the average number of
available upgrades per link. Note that these subnetworks have some overlap at the links
88 5. Case Study for the German Railway Network

leading from one subnetwork to another as well as the links connecting aggregated federal
states.

Table 5.2 summarizes the properties of all the instances with respect to network size,
demand gures and the available budget.

Network #Nodes #Arcs #Relations Annual budget [million e]


SH-HA 62 200 341 {6, 12, 18}
B-BB-MV 185 592 520 {1, 2, 3}
NS-BR 214 720 683 {90, 180, 270}
HE-RP-SL 251 814 527 {75, 150, 225}
TH-S-SA 246 778 672 {75, 150, 225}
BA-BW 369 1180 1119 {65, 130, 195}
NRW 389 1288 1133 {25, 50, 75}
Deutschland 1620 5162 3582 700

Table 5.2.: Properties of the test instances for our railway network expansion models

The instances visibly dier in size, which allows to investigate the scalability of the ap-
proaches. Furthermore, we consider three budget scenarios  low, medium and high 
for each subinstance, as this parameter has a signicant impact on the complexity of the
problem. The rates for revenue and transportation cost are the same as for the complete
instance. Planning and observation horizon shall always be 20 years respectively.

In the following, we analyse the computational properties of the models devised in this
work as well as the eects of the model improvements. In this, we restrict ourselves to
the seven subinstances of the German railway network with dierent budget choices, as
the entire Germany instance is too large to be processed in the memory of the employed
workstation, even after extensive preprocessing. The case study for this complete problem
instance, which is the actual subject of investigation, will be presented in Section 5.3 with
a comprehensive assessment of the developed solution algorithm.

5.2. Evaluating the Models and their Enhancements


All the considerations in Chapter 4 persue the aim to get the huge complexity of the
problem under control. They include the choice of the best model, strengthening model
formulation and eective preprocessing. Their eects are investigated in the following.

5.2.1. Comparing the Three Model Formulations for (MNEP)


Chapter 3 saw the introduction of two basic models for the multi-period network expansion
problem, (FMNEP) and (BMNEP). In Section 4.1, we were able to show that the LP
relaxation of (FMNEP) is always at least as good as that of (BMNEP). Furthermore, in
5.2 Evaluating the Models and their Enhancements 89

Section 4.2, we derived a more compact formulation of (FMNEP), which we denoted by


(CFMNEP). In the following, we give a computational comparison of the three models and
examine if it matches the theoretical ndings.

We set the time limit to one hour for each model and each instance and set the optimality
gap tolerance to zero. Table 5.3 shows the results.

Solution time / #BNB-Nodes or Gap


Instance (BMNEP) (FMNEP) (CFMNEP)

SH-HA-6 1142 s / 3058 251 s / 59 220 s / 16


SH-HA-12 425 s / 719 223 s / 162 271 s / 76
SH-HA-18 358 s / 590 233 s / 118 178 s / 45
B-BB-MV-1 1378 s / 150 1158 s / 48 938 s / 5
B-BB-MV-2 1002 s / 39 914 s / 0 684 s / 0
B-BB-MV-3 1173 s / 6 1010 s / 42 739 s / 0
NS-BR-90 1.59 % 0.13 % 0.16 %
NS-BR-180 0.32 % < 0.01 % 0.01 %
NS-BR-270 < 0.01 % < 0.01 % < 0.01 %
HE-RP-SL-75 0.70 % 0.20 % 0.15 %
HE-RP-SL-150 0.04 % 0.02 % 0.01 %
HE-RP-SL-225 < 0.01 % < 0.01 % < 0.01 %
TH-S-SA-75 0.04 % 0.03 % 0.02 %
TH-S-SA-150 0.04 % 0.01 % 0.01 %
TH-S-SA-225 0.01 % < 0.01 % < 0.01 %
BA-BW-65 9749.70 % 9749.70 % 2.11 %
BA-BW-130 1.12 % 0.08 % 0.03 %
BA-BW-195 0.12 % 0.02 % 0.04 %
NRW-25 − − −
NRW-50 8988.55 % 0.50 % 8988.55 %
NRW-75 − 0.25 % 8988.61 %

Table 5.3.: Computational behaviour of the models concerning the solution time and the
required number of branch-and-bound-nodes or the remaining gap after 1 h

For each model, we see its solution time and the required number of branch-and-bound
nodes to solve an instance in case of optimal solution and the remaining optimality gap
if it hit the time limit. The naming of the instances is subnetwork-budget, such that, for
example, SH-HA-6 refers to the subnetwork including Schleswig-Holstein and Hamburg
with a budget of 6 million Euros per year. Note that zero for the number of branch-
and-bound nodes indicates solution at the root node, while a dash indicates that the root
relaxation could not be solved within the time limit.

We see that Model (FMNEP) is clearly superior to (BMNEP) as there is not a single
instance which is solved faster or up to a lower gap by the latter. For those instances
which could be solved within the time limit, (BMNEP) tends to require much more nodes
90 5. Case Study for the German Railway Network

to solve it. It is obvious that the better LP bound of (FMNEP) plays out favourably,
conrming Theorem 4.1.1 from a computational point of view. Tightening its formulation
in Model (CFMNEP) improves its behaviour at large, with the only pathological excep-
tion of the NRW network. Here, Model (FMNEP) is faster in nding a good solution via
Gurobi's heuristics by luck, as it takes almost the whole time to solve the root relaxation.
This is similar for instance BA-BW-65 with Models (BMNEP) and (FMNEP). Altogether,
it is obvious to continue the computational study based on Model (CFMNEP).

These rst results also show the impact of the budget parameter as all subinstances tend
to be easier the more budget is available. This is in accordance with reality as making
plans is always easier with more money at hand.

5.2.2. Eciency of the Preprocessing


In the following, we examine the eciency of the preprocessing procedures developed in
Section 4.3. They aim at reducing the size of the routing block of (CFMNEP), i.e. the
number of routing variables and ow conservations constraints, which make up the biggest
part of the model.

The results on the instances and with the same parameter settings as above is shown in
Table 5.4. The rst two columns show the problem size in number of constraints and
variables as well as the reduction in size by our preprocessing respectively. Note that these
gures represent the values after Gurobi's built-in preprocessing, which allows us to come
to a better judgement of the eect of our own preprocessing. Obviously, the values in
both columns are largely unaected by the variance in the budget. At least for our own
preprocessing, this is not suprising, as it focusses more on the continuous routing block
than on the integral upgrade decisions. What we see is that our preprocessing based on the
maximal-detour criterion is extremely eective in reducing the problem size. For all the
instances, the number of constraints and variables can be reduced by half, in some cases
even by two thirds. This becomes even more impressive when considering that Gurobi's
built-in preprocessing is only able to get rid of about one per cent of the constraints and
variables. It can clearly be concluded that Gurobi itself is not able to detect and exploit
the shortest-path structure of the objective function.

The immense reduction in problem size leads to a considerable improvement of the com-
putational result within 1 h. On the one hand, the solution times of the easier instances
decrease sharply. On the other hand, there is no instance left for which the remaining
optimality gap after 1 h is higher than 0.23 %. This means that the network expansion
problem on these German subnetworks can be seen as solved.

5.3. The Germany Case Study


In this section, we turn our focus to the actual problem to be solved: nding an optimal
network expansion plan for the entire German railway network. We already saw that the
drastic reduction in size by our ecient preprocessing methods made subnetworks for one
5.3 The Germany Case Study 91

#Constraints / #Variables Solution time or Gap


Instance no PP [million] reduction by PP [%] no PP with PP

SH-HA-6 0.31 / 1.04 50.93 / 54.61 220 s 143 s


SH-HA-12 0.31 / 1.04 50.90 / 54.61 271 s 124 s
SH-HA-18 0.31 / 1.04 50.89 / 54.61 178 s 102 s
B-BB-MV-1 1.38 / 4.42 56.67 / 59.14 938 s 438 s
B-BB-MV-2 1.38 / 4.42 56.59 / 59.11 684 s 226 s
B-BB-MV-3 1.39 / 4.42 56.67 / 59.09 739 s 198 s
NS-BR-90 2.05 / 6.87 48.27 / 50.35 0.16 % 0.23 %
NS-BR-180 2.05 / 6.87 48.26 / 50.35 0.01 % 0.01 %
NS-BR-270 2.05 / 6.87 48.26 / 50.35 < 0.01 % < 0.01 %
HE-RP-SL-75 2.06 / 6.63 62.69 / 64.81 0.15 % 0.02 %
HE-RP-SL-150 2.06 / 6.63 62.68 / 64.81 0.01 % < 0.01 %
HE-RP-SL-225 2.06 / 6.63 62.68 / 64.81 < 0.01 % 2657 s
TH-S-SA-75 2.28 / 7.19 55.05 / 56.96 0.02 % 0.02 %
TH-S-SA-150 2.28 / 7.19 55.04 / 56.95 0.01 % < 0.01 %
TH-S-SA-225 2.28 / 7.19 55.04 / 56.95 < 0.01 % < 0.01 %
BA-BW-65 5.73 / 18.34 63.52 / 64.18 2.11 % 0.02 %
BA-BW-130 5.73 / 18.34 63.51 / 64.17 0.03 % < 0.01 %
BA-BW-195 5.73 / 18.34 63.51 / 64.17 0.04 % < 0.01 %
NRW-25 6.10 / 20.11 57.91 / 59.23 − 0.09 %
NRW-50 6.10 / 20.11 57.91 / 59.23 8988.55 % 0.01 %
NRW-75 6.10 / 20.11 57.91 / 59.23 8988.61 % < 0.01 %

Table 5.4.: The eect of Preprocessing for Model (CFMNEP) with respect to problem size
as well as solution time or remaining gap after 1 h

to three federal states solvable by standard methods. The network graph of the Germany-
wide instance, however, is more than four times larger than that of the biggest subnetwork
we considered, and the number of demand relations is more than three times larger. Even
after preprocessing, the mixed-integer program consists of 16 million constraints and 49
million variables and quickly brings the compute server to its memory limit when trying
to solve only the LP relaxation.

For this reason, we developed the multiple-knapsack decomposition scheme ( MKD) from
Section 4.4.1. Recall that it works by decomposing the problem along the timescale to
allow for a seperate treatment of each year of the planning horizon. As we consider a
planning horizon of 20 years, the problem can mainly be divided into 1 multi-commodity
ow network design problem as well as 20 individual multi-commodity ow problems, each
the size of the whole German railway network.

Employing this decomposition scheme, we were able to obtain a solution for this instance
within 8 hours. Here, we would like to add that the 20 multi-commodity ow subproblems,
which account for the biggest part of the solution time, were solved sequentially. It is easily
possible to speed up the solution process by solving the subproblems in parallel. This would
92 5. Case Study for the German Railway Network

allow to obtain the same solution within about 20 minutes on a machine with 21 processor
cores (an additional one for the master problem), which is more than sucient to serve as
an ecient optimization tool for planners at GSV.

Preliminary computation experiments on the subnetwork topologies from before showed


that two minor modications of the scheme presented in Section 4.4.1 lead to signi-
cantly better solutions. The rst one consists in slightly reducing the budget allocated to
the single-period planning problem which determines the candidate upgrades. Instead of
granting it the whole amount of money available over the planning horizon (which would
be 20 times the annual budget) it turned out to improve the quality of the nal solution
to reduce this amount by 2 %. This resulted to be the best value according to our experi-
ments. The most probable explanation for this eect is that a few less important upgrades
are ruled out from the beginning. These upgrades could not be built anyway as spending
the whole budget and having all upgrades ready at once is an idealization of the situation
where the budget has to be spread over 20 years and upgrades are nanced over several
years. This introduces a slight ineciency in the temporal availability of the money such
that not all upgrades determined by the idealized model can actually be implemented.
Moreover, ruling out these less important upgrades may help to improve the estimation of
the protability of the remaining upgrades.

The second observation was that a signicant amount of computation time could be saved
by solving the single-period problem only up to an optimality gap of 0.0005 %. This led to
practically no loss in solution quality in the nal solution. Thus, all computational results
for MKD presented in the following were achieved with these two modications.
Before we discuss the structure and the properties of the obtained solution for the Germany-
wide instance, we rst undertake some eort to assess its quality.

5.3.1. Assessing the Quality of the Solution

The results presented in the following are intended to evaluate the quality of the solution
obtained by MKD for the complete German railway network. As an optimal solution of
Model (CFMNEP) for this instance showed to consume more resources as the employed
workstation provides, especially much more memory, we need alternative means to assess
the quality of our solution. To do so, we make use of the subnetwork topologies already used
in Section 5.2. Put together, they form a complete cover of the whole German network. We
compare the quality of the MKD solution on these subnetworks with the optimal solution
obtained via a standard approach. To allocate an appropriate annual budget to each of
these subnetworks, we distribute the entire budget of 700 million Euros according to its
distribution to the individual states in our MKD solution for the whole of Germany. In
other words: If our MKD solution distributes n % of its expenses to the links belonging
to a certain subnetwork, then we equip this subnetwork with a total budget equal to n %
of the German total budget. The annual budget of each subinstance is then obtained by
dividing this value by 20, as this is the length of the planning horizon.

Table 5.5 shows the resulting budget values for each of the instances.
5.3 The Germany Case Study 93

Annual budget Total budget Relative size


Network [million e] [million e] of the budget

SH-HA 12 240 medium


B-BB-MV 5 100 high
NS-BR 180 3600 medium
HE-RP-SL 295 5900 high
TH-S-SA 60 1200 low
BA-BW 165 3300 medium-high
NRW 90 1800 high
Deutschland 700 14000 high

Table 5.5.: The budgets allocated to each instance

The sum over all annual budgets of the subinstances is somewhat higher than the 700
million Euros of the complete instance, which is due to the mentioned overlaps of the
subnetworks. We observed in preliminary experiments that the quality of the solutions
obtained via MKD generally improves with a higher annual budget given to an instance.
Therefore, we also state the relative size of the budget for each subinstance compared to the
budgets established in Section 5.1.2 to qualify the computational diculty of the problem.
From Table 5.5, we see that our set of reference instances for Method MKD covers the
whole range of low budgets (less money than needed for an optimal single-period solution)
over medium budgets (just enough money) up to high budgets (more money than needed).
For a discussion of the budget of the complete instance, we refer to Section 5.3.2.

The results of Algorithm MKD on these reference instances compared to an exact solution
are given in Table 5.6.

Solution time MKD quality


Instance (CFMNEP) MKD plain adjusted

SH-HA-12 124 s 10 s 0.06 % 1.69 %


B-BB-MV-5 145 s 56 s < 0.01 % 0.19 %
NS-BR-180 > 13605 s 301 s 0.09 % 1.19 %
HE-RP-SL-295 2158 s 163 s 0.02 % 0.22 %
TH-S-SA-60 24337 s 208 s 0.01 % 1.30 %
BA-BW-165 6595 s 685 s 0.04 % 0.48 %
NRW-90 > 7885 s 1016 s 0.06 % 1.32 %
Deutschland-700 − ≈8h 0.13 % < 0.78 %

Table 5.6.: Performance of multiple-knapsack decomposition  comparison of the solution


time against Model (CFMNEP) solved by Gurobi and quality of the solution by MKD
We compare the time needed for an exact solution of Model (CFMNEP) via Gurobi with
the execution time of MKD as well as the quality of the MKD solution. First, we observe
94 5. Case Study for the German Railway Network

that the heuristic solution is available in much shorter time than the exact solution for any
of the instances. Exact solution takes more than a factor of 10 longer in many cases, and
instances NS-BR-180 and NRW-90 could not be solved at all because exact solution went
out of memory. Again, we note that the 20 MKD subproblems were solved sequentially.
Thus, a signicant further time saving is possible by solving them in parallel.

To quantify the quality of the MKD solutions, we look at two dierent measures. The
rst one, indicated by plain in Table 5.6, states the deviation of the heuristic solution
from the optimal objective function value in per cent. In case Model (CFMNEP) could
not be solved to optimality, the resulting upper bound on the error is given. We see that
the heuristic solution never deviates more than 0.1 % from the optimal solution on any of
the subinstances, which practically means that it always gets the rst two digits after the
comma correctly.

Now, we have to take into account that the biggest part of the prot resulting from the
transportation of the demand is already achievable without implementing any upgrade at
all. This is because the biggest part of the network is already in place from the start.
Thus, what we are actually optimizing is the increase in prot achievable by the upgrades
compared to the prot possible without upgrading the network. This is why Table 5.6
also states a second measure of quality, adjusted solution quality, which is the actually
important one. It compares the increase in prot in the MKD solution to that in the
optimal solution. We see that MKD performs very well with respect to this measure. The
error in the adjusted objective function is clearly below 2 % for any of the instances, which
is more than satisfying for this real-world application.

We also see the tendency that higher annual budgets (cf. Table 5.5) make it easier for MKD
to nd good solutions. In fact, for three out of four subinstances with higher budgets, the
error is even below 0.5 %. Furthermore, we see no strong indication that the error increases
with the size of the instance.

The convincing results on subinstances allowed us to expect that our MKD solution for
the whole German railway network is of very high quality, too. The fact that this complete
instance is composed of the subnetworks under consideration in Table 5.6 contributes to
this expectation. Both the topology and the demand structure of the subinstances are
representative of the complete instance. This also applies to the regional distribution of
the budget as discussed above.

This is what we can say using the workstation described at the beginning of the chapter.
Later computations on a much stronger workstation allowed us to derive upper bounds
for an optimal solution to the problem. This way, we were able to show that our solution
for the whole of Germany is at most 0.78 % away from an optimal solution (adjusted),
which is a remarkable result for an instance of that dimension. Last but not least, the high
quality of the solution is also conrmed by planners at GSV, which we are going to detail
in the following section.

Altogether, we can say that MKD is a method which attains very satisfying solutions
within a short amount of time and with limited resources. This makes it well suited for
quick evaluations of dierent network expansions under varying demand scenarios, where
its small error is entirely negligible. Moreover, we remind that the extension of MKD
5.3 The Germany Case Study 95

to an exact algorithm presented in Section 4.4.2 is able to produce results of even higher
precision by iteratively improving the solution, if this is needed.

5.3.2. Presentation and Discussion of the Solution


In this section, we present our solution for the network expansion problem for the German
railway network given to us by our partners at GSV. To be able to compare the solution
obtained by our models and algorithms with the expected future demand situation, we
begin with an explanation of the target demand pattern for 2030. Then we examine our
solution from two dierent perspectives, namely the schedule of the expansion plan and
the distribution of the measures under consideration. Finally, we discuss general properties
of the solution and incorporate the point of view of our industry partner.

The Target Demand Scenario Figure 5.1 shows the growth in demand between 2010, the
base year of our study, and the target year 2030 according to the internal demand forecast
by GSV. Note that the gure is a joint presentation for freight trac and long-distance
passenger trac.

(a) Joint growth in freight trac and long- (b) Expected bottlenecks in the German rail-
distance passenger trac on the main trans- way network in 2030 without the creation of
portation corridors between 2010 and 2030 new capacities on the links

Figure 5.1.: Visualization of the GSV forecast for the growth in transportation demand
until the year 2030; source: DB Netz AG (2013), see also Beck (2013)

In Figure 5.1a, we see a map of the German railway network, where the links are coloured
in grey. A thicker line represents a higher utilization of the link in the base year 2010.
96 5. Case Study for the German Railway Network

The expected utilization of the links in the target year 2030 is superimposed in green, with
a thicker line representing a higher growth in demand. It is clearly visible that there a
certain corridors on which GSV expects a high increase in transportation demand. The
two most prominent ones are the corridor which leads from the north-west of Bayern
(Würzburg) up to Hamburg and the corridor coming from Switzerland, along the west
of Baden-Württemberg, up to the north-west of Nordrhein-Westfalen and going to the
Netherlands. This also aects the short corridor passing by Frankfurt, which connects
these two main corridors. Steep increases in demand are also visible between München
and the border to Austria, from Würzburg to the border of Austria, between Mannheim
and München, between Nürnberg and Leipzig, from Hamburg over Dresden to the border
of the Czech Republic as well as in the greater regions of Berlin and the Ruhrgebiet.

This projected development is put into the context of the available link capacities in Fig-
ure 5.1b. It shows the regions in the network for which the GSV forecast actually predicts
bottleneck situations. The links are coloured in grey with an increasing thickness for a
higher amount of trac. Expected bottlenecks are encircled in orange. This second g-
ure shows that the projected growth in demand leads to critical bottlenecks on the two
main corridors mentioned above as well as in the regions around München and north of
Hamburg. They entail the risk of decreasing punctuality and the increase of denied route
reservations for railway transport companies. An eective satisfaction of the demand in
mobility is only possible by the construction of new links and the upgrade of existing links
(all three quotations from DB Netz AG in Beck (2013)). The GSV forecast is somewhat
more moderate in its estimation of the future railway trac compared to the study by the
Umweltbundesamt (2010) cited in Chapter 1 (cf. Figure 1.2, page 13). Nevertheless, it
projects an overall growth of rail freight trac in the order 50 % over the planning horizon,
or 2 % per year on average.

Taking a second look at Figure 5.1, we see that not all the links with strong increases
in utilization are in danger of becoming bottlenecks. Especially in the eastern part of
Germany, current capacities seem to be mostly sucient to accommodate the increasing
trac. Here, it becomes obvious that an ecient expansion plan has to reect a trade-o
between rerouting part of the trac (possibly o its optimal route) and investing into new
capacities. Incorporating this trade-o is exactly what our model (CFMNEP) does. In
the following, we present its solution for an optimal expansion of the German rail freight
network between 2010 and 2030.

The Schedule of the Expansion Plan Figure 5.2 features a series of four pictures which
show the progressing expansion of the German rail network according to the solution
via multiple-knapsack decomposition. These pictures show the links that were chosen to
upgrade over the planning horizon, highlighting all realized upgrades until 2015, 2020, 2025
and nally the year 2030.
Comparing our solution with the demand forecast of Figure 5.1a, we see that it seems to
be a very adequate response to the growth in trac. Upgrades focus on the two main
corridors identied before as well as the short corridor connecting them, which are exactly
the routes with the highest increase in demand. Further upgrades focus on the regions of
München and Berlin, among others. We may say that our solution eectively addresses
5.3 The Germany Case Study 97

(a) Upgraded links until 2015 (b) Upgraded links until 2020

(c) Upgraded links until 2025 (d) Upgraded links until 2030

Figure 5.2.: The evolution of the German railway network between 2010 and 2030 under
the solution determined via multiple-knapsack decomposition  Each picture introduces a
new colour for the upgrades nished up to the corresponding year. The darker the colour
is, the later is the time of implementation.
98 5. Case Study for the German Railway Network

all the bottlenecks identied in Figure 5.1b. A slight exception is the segment north of
Hamburg, which was not part of our input data. Going beyond the pure necessity of
eliminating the bottlenecks, our solution relies on a general consolidation of the important
corridors to increase the eciency of the network. This is supported by the structure of
the objective function of Model (CFMNEP), which rewards taking the shortest path and
sets limits on possible detours. This criterion proves to be very well suited to focus on the
eciency of the most important connections in the railway network.

Following the series of pictures in Figure 5.2, it becomes visible that not all the upgrades
posses the same degree of urgency. It may be said that the bottlenecks from Figure 5.1b
tend to be resolved earlier in the planning horizon, while the general strengthening of the
important connections is conducted towards the end of the planning horizon. Obviously,
we cannot provide all the required capacity at once due to budget limitations, which makes
our gradual multi-period planning model a suitable approach. The consequence is that a
signicant number of links is revisited, i.e. subsequent upgrades on the same links are
performed later in the planning horizon. Our approach is all the more suited when taking
into account that network planning is a continuous process. Using our model, we can
identify the most important upgrades within the next ve years together with an outlook
into the future  an approach which can be repeated every ve years when updating the
network development plan.

Distribution of the Measures under Consideration In Figure 5.3, we give a second


perspective on our network expansion plan until 2030. It shows the nal layout out of the
network in 2030, where the upgrades are distinguished by the dierent kinds of underlying
measures.

It is visible at rst sight that many links need more than one kind of upgrade to attain
the desired level of capacity. This is particularly the case on the two main corridors in
north-south direction in the west and in the center of Germany.

The second striking feature of the expansion plan is the frequent use of block size reductions.
This is easily explainable by the properties of this measure. Block size reductions are cheap
and quick to implement and at the same time, they yield a moderate but yet non-trivial
increase in capacity (see Table 2.1, page 27). Thus, it is a measure which is well suited
as a quick remedy to urgent bottlenecks in the network. As our study is focussed on
the four most important measures to extend link capacities, the frequent use of block
size reductions may be interpreted as a general need of a link to quickly create some extra
capacity; a result which may also be achieved by other small-scale measures implementable
with limited eort and resources. Together with the rst observation of multiple upgrades
on the same links, it can be concluded that many links are in need of rst aid by short-term
upgrades until they obtain a bigger upgrade later in the planning horizon.

A third point to mention are the numerous upgrades of very short links, many of them even
below one kilometre. These upgrade proposals result from the fact that we did not impose
a restriction on the minimal length of a link to be upgraded. Of course, it is not realistic
to implement a massive upgrade on a link which only leads from a certain station to the
next junction. But the upgrades proposed for such links allow for another interpretation.
5.3 The Germany Case Study 99

Figure 5.3.: The upgrades chosen by multiple-knapsack decomposition up to the nal year
2030 distinguished by type of measure  Links marked in green are upgraded via block size
reductions, purple stands for the construction of new rails, blue indicates speed improve-
ments and red highlights electrications.
100 5. Case Study for the German Railway Network

They may be seen as the need to increase the throughput of the corresponding stations,
i.e. the number of trains which can be processed in such a station within a given amount
of time. Such node upgrades were not taken into account directly in this study, but in the
form of link upgrades they give an indication which stations are overloaded and as such
form a bottleneck in the network.

This second perspective on the obtained solution again conrms the suitability of the
model, as it allows a dierentiated view on the amount of capacity to install on each link.
In particular, it allows to upgrade a link more than once within the planning horizon to
install the capacity as needed.

General Discussion of the Solution Including both multiple upgrades on the same link
as well as those on short links emanating from a station, our solution proposes a network
expansion composed of about 600 upgrades of dierent scale within the planning horizon
of 20 years.

The budget of 700 million Euros per year is not entirely used up; about 10 % of the money
obtained are left after 2030. This justies the classication of the budget of the Germany
instance as high in Table 5.5. We already pointed out that a higher budget facilitates
the task of our multiple-knapsack decomposition algorithm to nd high-quality solutions.
From the point of view of the application, a higher annual budget means more exibility
in the design of the expansion plan and allows to implement urgently needed upgrades
earlier. On the other hand, the upgrades chosen in our solution are still widely spread over
the whole planning horizon. Thus, we can say that the annual budget of 700 million Euros
set in accordance with GSV was not calculated too high. In computational experiments,
we also tested a budget of 650 million Euros per year. This amount of money is sucient
to implement almost all the upgrades included in the solution presented above. The main
dierence in the outcome was that some of them were shifted to later years in the planning
horizon. Altogether, this indicates that the solutions proposed by our approach are quite
robust under a moderate reduction of the budget.

An important question to be considered is: Does the proposed upgrade eliminate all the
bottlenecks in the network until 2030? The answer has to be seen in a dierentiated
manner. About 2.5 % of the demand relations still cannot be routed at the end of the
planning horizon, at least not without taking unprotable detours. On the other hand,
as just discussed, the disposable budget is not completely used up. This implies that all
upgrades available in the input data which are able to contribute to a higher throughput of
the network have been implemented. The fact that several links exhaust all their upgrade
possibilities supports this nding. The same applies to the high number of articial node
upgrades which are part of the solution. Given such a nding, planners at GSV can now
use our software tool to reconsider the set of measures put into play and to evaluate an
increased spectrum of action. This way, our methods can give valuable decision support
in the highly iterative process of network planning.

At this point, it is very interesting to compare the results of our study with the network
development plan actually put forward by DB Netz AG. In the same time frame in which
this thesis came into being, the DB planners had to develop a network expansion strategy
5.3 The Germany Case Study 101

for the new Bundesverkehrswegeplan, which is due for 2015. Their expansion plan is shown
in Figure 5.4.

Figure 5.4.: The network upgrade devised by DB Netz AG for the Bundesverkehrswegeplan
2015  Upgrades are marked in red. Source: DB Netz AG (2013)

Several remarks have to be made to explain the dierence to our own results. Firstly, this
picture omits those upgrades that had already been initiated before 2010. They are con-
sidered as already implemented in Figure 5.4. In our study, they have been reevaluated. A
noticeable example is the expansion of the subcorridor Basel-Freiburg-Oenburg-Karlsruhe
which is part of the network plan by DB Netz AG and which has been conrmed to be
sensible in our study (cf. Figures 5.2 and 5.3). Secondly, their network plan includes the
construction of several completely new links, a measure that has not been considered here.
Finally, the most important point is the input data concerning the available upgrades for
the existing links. The most prominent example here is our proposal for an expansion
along a very long corridor through the center of Germany. This is an upgrade which is
very expensive to realize due to the landscape prole. The existing links here lead through
many tunnels, a feature that was not represented in our network data. The same applies
to increased costs due to bridges or high elevations. All the above points are easy to
incorporate into our models, however.

Despite the dierences mentioned above, planners at GSV are very satised with the
quality of the solution produced by our methods. They found that our modelling approach,
although undertaking a number of simplications, proved to be a suitable representation
of the real-world situation. Especially, they see the proposed upgrades as a plausible
reaction to the expected future demand patterns under the assumptions taken for this
study. According to them, the most important step to be undertaken is a more careful
compilation of the input data, such as more dierentiated cost factors for each link and
102 5. Case Study for the German Railway Network

some restrictions on the available upgrades, to increase the signicance of the results. Of
high importance to planners at GSV is the ability of our approach to generate solutions
which are not obvious at rst sight. This point came up in joint discussions as our plan
proposes the creation of a short bypass corridor through Rheinland-Pfalz to support the
longer corridor from Baden-Württemberg to Nordrhein-Westfalen. This is a feature not
present in their own network plan, which led to some surprise. An analysis of the underlying
trac ows can now tell them why the realization of this alternative route might be a viable
means to strengthen the western main corridor.

This is probably the outcome of such a real-world optimization project that should be
aspired: creating a solution that is not too far o of the expectations of the planners,
but also oering them new information to take into account in their work. Therefore, our
partners consider the developed software planning tool a valuable support for their network
planning in the future.
Part II.

Iterative Aggregation Procedures


for Network Design Problems
II. Iterative Aggregation Procedures for Network Design Problems 105

In modern applications, mathematical optimization problems usually have to be solved


for very large instances. As the problems are often NP-hard, this poses challenges to
state-of-the-art solution approaches. In particular, this is the case for network design
problems, where many real-world instances cannot be solved with the techniques that are
currently available. A question that arises naturally is whether the size of the problem
can be reduced to facilitate the solution. However, it is also desirable to maintain global
optimality in the sense that an optimal solution to the smaller problem can be translated
to a globally optimal solution to the original problem.

To this end, we develop an algorithm for network design problems based on aggregation,
which means the clustering of nodes in the underlying graph to components. This leads to
a coarser representation of the network with fewer nodes and arcs, which allows to solve
the corresponding network design problem in much shorter time. The global optimality
of the solution is ensured by embedding this approach into a disaggregation framework in
which the coarse representation of the network is adaptively rened where needed. The
algorithm terminates when the coarse solution is feasible for the original graph, in which
case we are able to prove that we have found an optimal solution.

The aggregation algorithm is rst derived for the simpler case of network design without
routing costs. We present three possible implementations of our framework, whose suitabil-
ity we examine for generic single- and multi-commodity network design problems. Most
importantly, this framework is extendable to more general problem settings. We demon-
strate this by showing how to incorporate costs for routing ow through the network by the
introduction of additional cutting planes. Our computational results on dierent bechmark
sets indicate the superior performance of our method compared to solving the problem via
a standard solver.

The structure of Part II is as follows. We begin with a description of the algorithmic


idea in Chapter 6, embedded into the context of the existing literature. This is followed
by a detailed derivation of our aggregation technique in Chapter 7, where we introduce
a basic version of the framework for network design problems without routing costs. It
features the decomposition into a master problem proposing network designs and a sub-
problem checking their feasibility. We will be able to prove that the rst feasible solution
we encounter in a sequential application of our approach is already an optimal solution.
Apart from this sequential realization and several enhancements for an ecient implemen-
tation, we propose two more versions of our method that allow for an integrated solution
of the problem within a branch-and-bound process. The advantage lies in retaining the
bound information after each disaggregation step as opposed to beginning solution from
scratch. In Chapter 8, we augment our framework by the additional cutting planes which
allow for the integration of routing costs. These cutting planes take the form of lifted
Benders optimality cuts. To motivate this approach, we elaborate on the relation of ag-
gregation and Benders decomposition in general and present a special case which permits
a combinatorial view on our idea. The theoretical derivation of our algorithms is rounded
o with extensive computational experiments in Chapter 9. Our benchmark sets include
randomly generated scale-free networks and real-world instances from Part I on the ex-
pansion of railway network. In both cases, we present results for both the basic version
of the aggregation framework and its extension integrating routing costs. They show the
106 II. Iterative Aggregation Procedures for Network Design Problems

superior performance of our algorithm compared to solving the problem via a standard
branch-and-bound implementation.
6. Motivation

Aggregation in the context of mixed-integer programming is a coarsening process that


omits details but ensures a global view on the complete problem. There are two common
reasons for performing aggregation of the problem data. Firstly, a problem instance may
be too large for solving it at its full size. Secondly, the data is imprecise, e.g. due to
measurement or forecasting errors, and thus a rough solution over an aggregated version
of the data is preferred over an expensive optimal solution. Disaggregation, on the other
hand, can be seen as an inverse procedure that reintroduces more detailed information.

Iterative aggregation and disaggregation techniques typically cluster parts of the original
problem and solve the arising aggregated instance. Then a certain update step is under-
taken, yielding a more or a less aggregated instance respectively. Alternatively, the degree
of aggregation may be preserved, but the aggregation rule is updated. This process is
iterated until some stopping criterion is satised.

Such techniques have successfully been applied to continuous linear programming problems.
Balas (1965) suggests a solution method for large-scale transportation problems that does
not consider all data simultaneously. Zipkin (1977) derives a posteriori and a priori bounds
for general linear programs. A comprehensive survey on aggregation techniques is given in
Dudkin et al. (1987), and Litvinchev and Tsurkov (2003) wrote a book about aggregation in
the context of large-scale optimization problems. In addition, aggregation techniques have
been applied to a wide eld of applications, e.g. network ow problems (Francis (1985),
Lee (1975)) and the optimization of production planning systems (Leisten (1995), Leisten
(1998)).

The use of aggregation techniques in mixed-integer problems is still not too widespread
but steadily expanding. Rosenberg (1974) describes aggregation of equation constraints,
while Chvátal and Hammer (1977) analyse the aggregation of inequalities, and Hallefjord
and Storoy (1990) use column aggregation in integer programs. Karwan and Rardin (1979)
investigates the relationship of Lagrangean relaxation to surrogate duality in integer pro-
gramming.

Aggregation has proved very useful for handling highly symmetric problems. In Linderoth
et al. (2009), it is one of several tools to grind a very hard problem instance from coding
theory. Salvagnin and Walsh (2012) use aggregation to form a master problem of a decom-
position method for multi-activity shift scheduling. Especially, shortest path algorithms
based on graph contractions (see Geisberger et al. (2012)) are very successful in practice.
Recent examples for the use of aggregation include Macedo et al. (2011) for a vehicle
routing application with time windows and Newman and Kuchta (2007) for scheduling
the excavation at an underground mine. Furthermore, Schlechte et al. (2011) present a
micro-macro transformation for railway networks that is used to solve the track allocation

© Springer Fachmedien Wiesbaden 2016


A. Bärmann, Solving Network Design Problems via Decomposition,
Aggregation and Approximation, DOI 10.1007/978-3-658-13913-1_7
108 6. Motivation

problem. An extensive case study on this approach is given in Borndörfer et al. (2014a).
In Borndörfer et al. (2014b), the authors introduce a column generation procedure with
a coarsened pricing-problem and successfully apply their method to optimal rolling-stock
planning in railway systems. Caimi (2009) proposes a two-level approach for the generation
of railway timetables which is based on a decomposition of the railway network into areas
with high trac density and areas with low trac density.

We are only aware of two previous applications of aggregation to network design problems.
These are Crainic et al. (2006), where a multi-level tabu search for the problem is developed,
and Boland et al. (2013), where iterative aggregation is applied to the planning horizon in
a network design problem with time-dependent demand.

The idea we present in the following is to solve network design problems by aggregation of
the underlying graph structure. Our aim is to develop an exact algorithm for the problem
that is based on an iterative renement of the network representation. An important
observation on the way to such an algorithm is that network design problems very often
possess optimal solutions which only include a small fraction of all available arcs. This is
illustrated by popular combinatorial optimization problems which can be stated as network
design problems, such as the minimum-spanning tree problem, the Steiner tree problem or
the travelling-salesman problem. But also more general network design problems typically
exhibit this property, which becomes even more frequent with an increasing degree of initial
capacities in the network.

In particular, it is very common in the case of infrastructure network expansion that


a relatively developed network has to be upgraded in order to allow for the routing of
additional demand requirements. In this situation, normally only a small percentage of
the arcs has to be upgraded. These arcs are frequently referred to as bottlenecks. Such
bottlenecks constitute the limiting factor for additional demand to be routed on top of
those that can already be accommodated. A striking example is given by the Germany-
wide instance of the railway network expansion problem considered in Part I. In 2010, the
initial year of our study, the German railway network was able to accommodate about 80 %
of the forecasted demand for the year 2030. An appropriate upgrade requires capacity-
increasing measures on less than 20 % of the links.

This fact was our motivation to devise an iterative algorithm based on aggregation that
continuously updates a set of potential bottleneck arcs to consider for the nal network
design. It works by clustering the nodes of the network graph to components. These
components are the nodes in a new, coarse representation of the network. The idea is to
choose the clustering such that the arcs connecting them are exactly the bottlenecks in the
network. For some initial aggregation of the network, we solve the corresponding network
design problem. Then we check if the arising aggregated solution can be extended to a
feasible solution for the original network. If so, we are able to prove that this extended
solution is optimal. If not, we get a certicate where to rene the aggregated representation
of the network and the problem is resolved.

Our method is rst presented for network design without routing costs. This is the simpler
case as we only have to ensure that the ow in the aggregated network induces a feasible
6. Motivation 109

ow within each component. We will see that the proposed algorithm results in a cutting-
plane procedure. By examining the relation of this procedure to Benders decomposition we
prove that our cutting planes strictly imply the corresponding Benders feasibility cuts.

Afterwards, we show that it is possible to incorporate the routing cost of the network ow
by introducing additional cutting planes. These are derived from the Benders optimality
cuts for a suitable decomposition into master problem and subproblem. The result is a
hybrid version of Benders decomposition, where the feasibility cuts are replaced by our
aggregation cuts. Moreover, this hybrid version permits to change the proportions of the
decomposition in the process, i.e. variables are allowed to move from the subproblem to
the master problem, which is not the case in an ordinary Benders decomposition.

Finally, we give computational results for our framework on two dierent benchmark sets.
These are random scale-free networks and instances derived from the real-world railway
networks considered in Part I. Results on further benchmark sets can be found in Bärmann
et al. (2015b). Together, they show the superior performance of our aggregation scheme
compared to solution of the problem via a standard solver. It turns out that the aggregated
graphs at the point where optimality is proved are usually much smaller than the original
graph, which results in a signicant speed-up. But even when we have to disaggregate
until reaching the original graph, our iterative approach mostly outperforms a solution of
the complete problem at once. Furthermore, we are able to show that our method can be
extended successfully to the case with routing costs.
7. An Iterative Aggregation Algorithm for
Optimal Network Design

In this chapter, we develop an exact algorithmic scheme for the solution of network design
problems which is based on iterative (dis)aggregation of the underlying network graph. For
the ease of exposition, we use a very basic problem setting for the derivation of our frame-
work, namely a canonical single-commodity network expansion problem without routing
costs and a single module per arc (cf. the hierarchy of network design problems given in
Section 2.3.1). Two remarks are important about this choice. Firstly, the problem un-
der consideration is already NP-complete (by specialization to the Steiner tree problem).
Secondly, the algorithmic framework developed for this basic problem can be generalized
to a much broader class of network design problems. Our results presented in Chapter 9
demonstrate that a straightforward extension to multi-commodity ow problems with sev-
eral modules per arc is possible. Moreover, we show in Chapter 8 that it is possible to
incorporate routing costs via additional cutting planes.

Slightly modifying the notation established in Section 2.3.1, the network expansion problem
in its single-commodity variant with a single module per arc can be stated as the following
mixed-integer problem:

P P
min k a ua + f a ya
a∈A a∈A
P P
s.t. ya − ya = d v (∀v ∈ V )
a∈δv+ a∈δv−
(7.1)
ya ≤ ca + Ca ua (∀a ∈ A)
ua ∈ Z+ (∀a ∈ A)
ya ≥ 0 (∀a ∈ A).

In the single-commodity version of the problem, we have a parameter dv ∈ for the R


demand of node v ∈ V . Furthermore, we remark that Problem (7.1) is formulated requiring
ya ≥ 0 instead of ya ∈ [0, 1], a ∈ A, in order to allow for a compact formulation with |A| ow
variables. In the other case, we would have to use the multi-commodity ow formulation
of the single-commodity problem using |A| · |R| variables, where R is the set of all ordered
pairs of nodes in v∈V with non-zero demand dv . Finally, we assume for the rest of this
chapter that the routing costs fa are equal to zero and drop them from the problem. That
means were interested in a least-cost capacity upgrade of the arcs to full the demand
requirements of the network.

The main algorithmic idea persued in the following is a decomposition of the problem
which  from a bird's eye view  can be stated as follows:

© Springer Fachmedien Wiesbaden 2016


A. Bärmann, Solving Network Design Problems via Decomposition,
Aggregation and Approximation, DOI 10.1007/978-3-658-13913-1_8
112 7. An Iterative Aggregation Algorithm for Optimal Network Design

1. Partition the node set of the graph into components, i.e. choose an initial aggregation.

2. Master problem: Solve the network expansion problem over the aggregated graph.

3. Subproblem: Check the feasibility of the network upgrade w.r.t. the original graph.

4. a) In case of feasibility: Terminate and return the feasible network expansion.

b) In case of infeasibility: Rene the partition and go to Step 2.

Our presentation of the method begins with the denition of the master and the subprob-
lem and a detailed explanation of its intermediate steps. We continue with a proof of its
correctness, i.e. we show that it always terminates in a nite number of iterations, return-
ing an optimal solution to the original problem. Finally, we introduce several possible
enhancements of the algorithm whose incorporation turns out to be very favourable in the
light of our computational results.

7.1. Graph Aggregation and the Denition of the Master


Problem
The term aggregation takes a variety of dierent meanings in the literature. In the context
of our algorithm for optimal network design, aggregating a directed graph G = (V, A)
means clustering its nodes into subsets. aggregated graph Gϕ = (Vϕ , Aϕ )
We dene the
with respect to a surjective clustering function ϕ : V → {1, . . . , k}, k ∈ N, as follows. Its
node set Vϕ = {V1 , . . . , Vk } is a partition of V into k components with v ∈ Vi ∈ Vϕ i
ϕ(v) = i ∈ {1, . . . , k}. Its arc set Aϕ contains a directed arc from Vi ∈ Vϕ to Vj ∈ Vϕ for
each arc (u, v) ∈ A with i = ϕ(u) 6= ϕ(v) = j , i.e. u and v belong to dierent components.
Note that G as well as Gϕ are allowed to contain multiple arcs between the same two nodes.
Figure 7.1 illustrates the above denitions. Its shows a graph G overlaid with a possible
aggregation Gϕ . Nodes in the same component are encircled. Only those arcs connecting
dierent components in Vϕ are part of Aϕ .

The master problem of our algorithm is Problem (7.1) applied to some aggregation Gϕ of
G
P, which is done as follows. The aggregated demand vector dϕ is dened via dϕ (Vi ) =
v∈Vi dv for all i = 1, . . . , k , i.e. the demand of a component is the net demand of the
nodes it contains. The capacity cϕ (a) and the installable module of an arc a ∈ Aϕ are
those of the corresponding original arc. In order to simplify the notation, we identify a
component Vi ∈ Vϕ with its index i and identify each arc a ∈ Aϕ with the corresponding
original one in A. The master problem with respect to Gϕ can then be stated as
P
min k a ua
a∈Aϕ
P P
s.t. ya − ya = d i (∀i ∈ Vϕ )
a∈δi+ a∈δi−
(7.2)
ya ≤ ca + Ca ua (∀a ∈ Aϕ )
ua ∈ Z+ (∀a ∈ Aϕ )
ya ≥ 0 (∀a ∈ Aϕ ).
7.1 Graph Aggregation and the Denition of the Master Problem 113

12

15

10

1
3
14
2

11

13
4
5

Figure 7.1.: A graph G = (V, A) and its aggregation Gϕ = (Vϕ , Aϕ ) w.r.t. the clustering
function ϕ with ϕ−1 (1) = {1, 7, 10}, ϕ−1 (2) = {2, 3, 4, 5, 6}, ϕ−1 (3) = {8, 9, 12, 15} and
ϕ−1 (4) = {11, 13, 14}.

By the denition of the aggregated demand vector, the ow conservation constraint for
component i ∈ Vϕ is exactly the sum of the original ow conservation constraints belonging
to the nodes v ∈ Vi . Furthermore, the capacity constraint for some arc a ∈ A of the original
problem (7.1) is part of the aggregated problem (7.2) if and only if a ∈ Aϕ holds. This
applies similarly to the summands in the objective function. Consequently, our master
problem is a relaxation of the original network expansion problem, and its optimal value
w.r.t. an arbitrary clustering function ϕ provides a lower bound for the optimal value of
the latter.

By construction of the aggregated problem, a feasible solution naturally translates to a (not


necessarily feasible) solution for the original problem by performing two steps: First, all
upgrade decisions (u-variables) corresponding to arcs within any of the components are set
to zero. Second, the induced ows (y -variables) within the components are computed by
solving a maximum-ow problem. This extendibility test is the purpose of the subproblem,
whose derivation is the content of the following section.

7.1.1. Denition of the Subproblem and Graph Disaggregation

The solution to Master Problem (7.2) induces new demands within the aggregated compo-
nents. The purpose of the subproblem is to validate whether these demands can be routed
114 7. An Iterative Aggregation Algorithm for Optimal Network Design

without additional capacity upgrades inside the components. This validation decomposes
into seperate subproblems for each component. An example of this situation is depicted
in Figure 7.2.

12
-4

0 15

10 9 6 3 -8 -4

15
9
-15 5
7

8 0
5 13 -8

6
1
3 3
14
2

15
6

11
-6
5

18
13
4
5

(a) Induced demands for component {8, 9, 12, 15} (b) The associated feasibility subproblem

Figure 7.2.: Illustration of the subproblem for some component of Gϕ

Figure 7.2a shows part of the solution to the master problem which induces a new demand
vector within some component of the aggregated graph. The corresponding subproblem
for this component is depicted in Figure 7.2b.

Checking the feasibility of a network expansion involves the solution of a maximum-ow


problem within each component as follows. Let Hi = (Vi , Ai ) be the subgraph of G = (V, A)
induced by component Vi of the partition of V according to ϕ. The nodes Vi of Hi have
an original demand of dv , v ∈ Vi . The optimal ows of the master problem induce new
demands within Hi as each ow ya on an arc a = (u, v) ∈ Aϕ with u ∈ Vi and v ∈ Vj
changes the demand of u to d ¯u := du + ya and that of v to d¯v := dv − ya . By introducing
articial nodes s as super source and t as super sink, the check whether a feasible ow
exists can be formulated as the following single-source maximum-ow problem:
P
max zv
v∈Vi+
v ∈ Vi+

zv ,
 if
v ∈ Vi−
P P
s.t. ya − ya = −zv , if (∀v ∈ Vi )
a∈δv+ a∈δv− 0,

otherwise
(7.3)
ya ≤ c a (∀a ∈ Aϕ )
zv ≤ |d¯v | (∀v ∈ Vi+ ∪ Vi− )
ya ≥ 0 (∀a ∈ Aϕ )
zv ≥ 0 (∀v ∈ Vi+ ∪ Vi− ).
7.1 Graph Aggregation and the Denition of the Master Problem 115

Here, Vi+ , Vi− ⊆ Vi are the nodes with positive (resp. negative) induced demand d¯v . Vari-
+
able zv for v ∈ Vi then models the ow from the super source s to node v , while zv for
v ∈ Vi− Pis the ow from node v to the super sink t. If the maximum s-t-ow attains a
¯
value of
v∈V + dv , the induced demands within component Vi are feasible, otherwise they
i
are infeasible.

A feasible component Vi requires no further examination in the current iteration. An


infeasible subproblem can occur for two reasons. It either suggests that the initial capacities
within the associated component are not sucient to route the demands induced by the
master problem solution. In this case, it was not justied to neglect the capacity limitations
within the component. Or the algorithm might not be able to prove that all upgrade
decisions are already optimal. We will explain in Section 7.2, paragraph Global Feasibility
Test, how to detect this case. Whenever an infeasible subproblem is encountered, the
partition is rened in order to consider additional arcs in the master problem. This arc
set is chosen as a minimum s-t-cut that limits the ow; see Figure 7.3a, which continues
the example from Figure 7.2b. The algorithm terminates as soon as all subproblems are
feasible.

12

15
-4 10

8
-8
6

1
3 3
14
2
15

-6 11

5
13
4

18 5

(a) Disaggregation along a minimal cut (b) Resulting graph after disaggregation

Figure 7.3.: Disaggregation of an infeasible component

Updating the master problem (7.1) is done by disaggregating the infeasible component
along this minimum cut as shown in Figure 7.3a. Without loss of generality, let Vk be an
infeasible component and let Vk1 , . . . , Vkl be the components into which Vk disaggregates.
We dene a new clustering function ϕ̄ : V → {1, . . . , k + l − 1} with ϕ̄(v) = ϕ(v) if v ∈
/ Vk
i
and ϕ̄(v) = k + i if v ∈ Vk ⊂ Vk ; see Figure 7.3b. In the next iteration of our algorithm,
we solve the master problem problem for the resulting aggregated graph Gϕ̄ .
116 7. An Iterative Aggregation Algorithm for Optimal Network Design

7.1.2. Correctness of the Algorithm


We show next that for zero routing costs and non-negative expansion costs, the above
method always terminates with an optimal solution to the original network expansion
instance.

Theorem 7.1.1 (Bärmann et al. (2015b)). For f = 0 and k ≥ 0 in Problem (7.1),


the proposed algorithmic scheme always terminates after a nite number of iterations and
returns an optimal solution to the network expansion problem for the original graph.

Proof. Termination follows from the fact that only nitely many disaggregation steps are
possible until the original graph is reached. Clearly, the returned solution is feasible for
the original network by the termination criterion. In order to prove optimality, let (ū, ȳ)
be the optimal solution to the nal master problem that has been extended to a solution
of the original problem as described above. Furthermore, let (û, ŷ) be an arbitrary solution
to the network expansion problem for the original graph G. We split the expansion costs
of the two solutions into the cost kM arising at the arcs Aϕ of the nal aggregated graph
Gϕ and the cost kS arising at the arcs A \ Aϕ . The corresponding expansion variables are
uM and uS . We derive
k T ū = kM
T
ūM + kST ūS = kM
T
ūM
as no expansion is performed within the components. Because ūM belongs to an optimal
upgrade of Gϕ , we can conclude
T T
kM ūM ≤ kM ûM .
The non-negativity of the expansion costs k then leads to

T T
kM ûM ≤ kM ûM + kST ûS = k T û.

Altogether, we have shown k T ū ≤ k T û and therefore (ū, ȳ) is an optimal solution to the
original problem.

Theorem 7.1.1 states that the proposed method is an exact algorithm for Problem (7.1).
We remark that the approach also gives rise to a potentially promising heuristic by taking
any of the intermediate master solutions and making them feasible via additional upgrades
within the components.

7.1.3. Relation to Benders Decomposition


The aggregation procedure developed in this chapter possesses some obvious similarities to
Benders decomposition. Both algorithms solve a succession of increasingly tighter relax-
ations of the original problem, which is achieved by introducing cutting planes. In the case
of the aggregation framework, these are part of the primal (aggregated) ow conservation
and capacity constraints. For Benders decomposition, these are the Benders feasibility and
optimality cuts. Both algorithms stop as soon as the optimality of the relaxed solution is
proved. Furthermore, the subproblem used in the aggregative approach coincides with the
7.1 Graph Aggregation and the Denition of the Master Problem 117

subproblem in Benders decomposition if the y -variables for the arcs in Aϕ are chosen to
belong to the Benders master problem.

However, there are also substantial dierences. The aggregation scheme introduces both
new variables and constraints in each iteration to tighten the master problem formulation.
Contrary to this, Benders decomposition is a pure row-generation scheme. Equally impor-
tant is the fact that the continuous disaggregation of the network graph leads to a shift
in the proportions between the master problem and the subproblem. The master problem
grows in size, while the subproblem tends to shrink as bottleneck arcs are transferred from
inside the components to the master graph. In comparison, Benders decomposition leaves
these proportions xed.

The following therorem details the relation between the subproblem information used in
the two algorithms.

Theorem 7.1.2 (Bärmann et al. (2015b)). Let ϕ be a clustering function according to a


given network graph G. For disaggregation of G along a minimal cut, the primal constraints
introduced to the master problem (7.1) in the proposed aggregation scheme strictly imply
the Benders feasibility cut obtained from the corresponding subproblem (7.3).

Proof. We prove the claim for the special case where the whole graph is aggregated to a
single component, i.e. ϕ≡1 and thus Aϕ = ∅. The corresponding situation in Benders
decomposition is that all arc ow variables are projected out of the master problem. The
extension of the arguments to the general case is straightforward.

In both algorithms, the task of the subproblem consists in nding a feasible ow in the
network and thus a solution to the following feasibility problem:

min 0
P P
s.t. ya − ya = d v (∀v ∈ V )
a∈δv+ a∈δv−
ya ≤ ca + Ca ūa (∀a ∈ A)
ya ≥ 0 (∀a ∈ A),

where ū corresponds to the network design solution determined by the master problem.
Let α denote the dual variables of the ow conservation constraints and β those of the
capacity constraints. In case of infeasibility, Benders decomposition derives its feasibility
cut from an unbounded ray of the dual subproblem:
P P
max dv α v − (ca + Ca ūa )βa
v∈V a∈A

s.t. αv − αw − βa ≤ 0 (∀a = (v, w) ∈ A)


βa ≥ 0 (∀a ∈ A).

Variables β can be eliminated from the problem, as any feasible solution (α̂, β̂) is dominated
by the solution (ᾱ, β̄) with ᾱ := α̂ and

β̄a := max{α̂v − α̂w , 0}


118 7. An Iterative Aggregation Algorithm for Optimal Network Design

for a = (v, w) ∈ A. Note that the dual subproblem is exactly that of determining a minimal
cut in the network. Thus, an arc a∈A is part of the minimal cut if and only if β̄a > 0,
which is exactly the case when the ow conservation constraints corresponding to the two
end nodes of a obtain dierent dual values. In case of an infeasible (primal) subproblem,
Benders feasibility cut can now be written as

X X
ᾱv dv ≤ max{ᾱv − ᾱw , 0}(ca + Ca ua ),
v∈V a=(v,w)∈A

for ᾱ belonging to an unbounded dual ray (ᾱ, β̄). For the same infeasible master solution,
the aggregation scheme adds the following system of inequalities to its master problem:

X X
ya − ya = d i (∀Vi ∈ Vϕ̄ )
+ −
a∈δV a∈δV
i i

and

ya ≤ c a + C a u a (∀a ∈ Aϕ̄ )
ya for a ∈ Aϕ̄ , where ϕ̄ is the clustering function induced by
together with the new variables
the minimal cut. As stated above, the dual subproblem values ᾱv and ᾱw coincide if nodes v
and w belong to the same component Vi ∈ Vϕ . Thus, we can dene αi := ᾱv for some v ∈ Vi .
The claim then follows by taking the sum of the aggregated ow conservation constraints
weighted with ᾱi and the aggregated capacity constraints weigthed with − max{ᾱv −ᾱw , 0}.
For this weighting, the left-hand side becomes

 
X X X  X
ᾱi  ya − ya  − max{ᾱv − ᾱw , 0}ya ,
Vi ∈Vϕ + − a=(v,w)∈Aϕ̄
a∈δV a∈δV
i i

which can be transformed to

X
(ᾱv − ᾱw − max{ᾱv − ᾱw , 0})ya .
a=(v,w)∈Aϕ̄

This yields
X
(ᾱv − ᾱw )ya ,
a=(v,w)∈Aϕ̄ :
ᾱw ≥ᾱv

which is non-positive due to the non-negativity of y. Thus, we nd

X X X X
ᾱv dv = ᾱi di ≤ (ᾱv − ᾱw )ya + max{ᾱv − ᾱw , 0}(ca + Ca ua ),
v∈V Vi ∈Vϕ a=(v,w)∈Aϕ̄ : a=(v,w)∈A
ᾱw ≥ᾱv

which implies Benders feasibility cut. Finally, this inequality is strict for all solutions to
the problem where ow is sent both along a certain arc of the minimal cut as well as along
its opposite arc. This completes the proof.
7.2 Possible Enhancements of the Algorithm 119

Theorem 7.1.2 shows that each single iteration of the aggregation scheme introduces more
information to the master problem than a Benders iteration. Although Benders decompo-
sition is often used to solve network design problems, it belongs to the folklore of mathe-
matical programming that the original Benders cutting planes are weak and numerically
instable already for small-scale networks. And this becomes even more problematic in the
case of large-scale networks as they are considered here. Therefore, Benders decomposition
is most commonly employed for smaller network problems with complicating constraints,
such as they arise in the context of demand stochasticity.

7.2. Possible Enhancements of the Algorithm


In its basic version, our algorithm is an iterative approach in which master problem and
subproblem are solved in an alternating fashion. This has the obvious drawback that
only very limited information is retained when moving from one iteration to the next. To
overcome this problem, it is possible to integrate the aggregation scheme into a branch-
and-bound framework. Furthermore, it is possible to use a hybrid of both approaches. In
our computational results in Chapter 9, we study the performance of all three variants.
The details of the resulting algorithms are given next.

Sequential Aggregation The Sequential Aggregation Algorithm ( SAGG) proceeds in a


strictly sequential manner: In each iteration, the network expansion master problem is
solved to optimality. In case of feasibility of all subproblems, the algorithm terminates.
Otherwise, the graph is disaggregated as described in the previous section, see Figure 7.4a
for a schematic example. For speeding up the rst iterations, we solve the network expan-
sion linear programming (LP) relaxations only and disaggregate according to the obtained
optimal solutions. Our experiments suggest that the savings in solution time compensate
for the potentially misleading rst disaggregation decisions based on the LP relaxation.
Note that the proof of Theorem 7.1.1 does not require u to be integral.

Integrated Aggregation In SAGG, most information  including bounds and cutting


planes  is lost when proceeding from one iteration to the next. Only the disaggregation is
carried over to the next iteration. The idea behind the Integrated Aggregation Algorithm
IAGG) is to retain more information on the solution process by embedding the disag-
(
gregation steps into a branch-and-bound tree. We start with some initial aggregation and
form the corresponding master problem. Each integral solution (x, y) found during the
branch-and-bound search is immediately tested for extendibility. In case of feasibility, we
allow (x, y) to become the next incumbent solution, otherwise we disaggregate the graph
and refuse (x, y); see Figure 7.4b for a schematic example. As noted in Section 7.1, aggre-
gation of a network can be performed by removing some constraints and adding up some
others from the formulation of the original instance. All constraints with respect to an
aggregated graph remain valid after disaggregation, which amounts to inserting the ow
conservation constraints for the new components and the arc capacity constraints for the
120 7. An Iterative Aggregation Algorithm for Optimal Network Design

arcs entering the master problem to the problem formulation. Altogether, this leads to a
cutting-plane algorithm for the network expansion problem.

The Hybrid Aggregation Algorithm A natural composition of SAGG and IAGG is the
Hybrid Aggregation Algorithm ( HAGG). It starts with a number of sequential iterations
such as in SAGG and then switches to the integrated scheme as shown in Figure 7.4c.
The idea is to have more information about the graph available at the root node of the
branch-and-bound tree once we start to employ the integrated scheme. This is favorable for
the cutting planes generated at the root node. For the rst iterations, HAGG and SAGG
behave exactly the same. Namely, we solve the LP relaxation of the master problem and
proceed with the obtained fractional solution. As a heuristic rule, we switch from the
sequential to the integrated scheme when the value of the LP relaxation of the master
problem is equal to the value of the LP relaxation of the original problem. Then the
optimal fractional expansion is found and the LP relaxation does not give any further
information for disaggregation.

(a) The Sequential (b) The Integrated Algorithm (IAGG); (c) The Hybrid Algorithm (HAGG);
Algorithm (SAGG) solution within a branch-and-bound tree using an advanced initial aggregation

Figure 7.4.: Schematic outline of the three aggregation schemes

Figure 7.4 summarizes the characteristics of the three aggregation schemes as described
above. Nodes labeled Int correspond to feasible integral solutions of the current master
problem, whereas Int / Inf indicates that such a solution is infeasible for the original
problem, leading to disaggregation. For the branch-and-bound trees, white nodes labeled
Inf or Frac represent infeasible and fractional branch-and-bound nodes (at which branching
might occur) respectively.
7.2 Possible Enhancements of the Algorithm 121

Apart from these global decisions on the implementation of our method, there are several
smaller algorithmic considerations, which are discussed in the following.

The Initial Aggregation The easiest choice is to aggregate the whole graph to a single
vertex so that the disaggregation is completely determined by the minimum-cut strategy.
As indicated above, this might not be the most suitable choice for the integrated scheme,
which is the motivation for a hybrid scheme. In fact, one can view HAGG as a version of
IAGG with a special heuristic for nding the initial aggregation.

Solving the Subproblems The subproblems are maximum-ow problems that we solve
via a standard LP solver, which is fast in practice. A potential speed-up by using specialized
maximum-ow implementations is negligible as the implementation spends almost all of
the time in the master problems and not in the subproblems.

Disaggregation We have seen how disaggregation works for a single component in Sec-
tion 7.1.1. However, in case of several infeasible components, it is not clear beforehand
whether all of them should be disaggregated, or otherwise, which one(s) should be used for
disaggregation. In our implementation, we always disaggregate all infeasible components,
which aims at minimizing the number of iterations. Experiments with other disaggregation
policies did not lead to considerable improvement. In addition, we also split components
that are not arc-connected into their connected components.

Global Feasibility Test It should be mentioned that the current network expansion might
be optimal although the extendibility test fails for some component. This is because not
only the expansions (u-variables) are xed in the subproblem but also the ow on all
edges contained in the master problem. To overcome this problem we can use a simple
global feasibility test: We x the expansion and check feasibility of the resulting maximum-
ow problem on the complete graph. The case where this global test saves at least one
iteration happened regularly in our experiments. As each iteration is relatively expensive,
this feature should denitely be included in the default settings. Note that the global test
cannot replace the subproblems for the components entirely as the cut it implies might
only consist of edges that have already been added to the master problem. Therefore,
after nding a solution to the master problem, we rst apply the global test and, in case
of infeasibility, we use the component subproblems to determine how to disaggregate. As
the global test only consists of solving one maximum-ow problem, it is performed in each
iteration.
8. Integration of Routing Costs into the
Aggregation Scheme

The preceding chapter saw the development of an aggregation-based algorithmic scheme


to solve large-scale network design problems. Its main idea is the aggregation of single
nodes of the network graph to components. Apart from reducing the size of the network
design problems, it also features a reduction of (continuous) symmetry in the problem
by neglecting the routing decisions within the components. This is possible as long as
these routing decisions are not valued in the objective function, i.e. when focussing on the
costs for installing new arc capacities only. In the following, we drop this assumption and
consider the general network expansion problem (7.1) with non-negative routing costs f.

The motivation to do so is given by many practical problem settings in which the network
routing is crucial to obtain an ecient solution. For example, this is the case with the
railway network expansion problem considered in Part I. Here, the optimal routing pattern
belonging to each possible network expansion decides whether the latter enables a more
protable network usage or not. For this reason, we will enhance the aggregation scheme
developed so far such that is able to incorporate the routing costs arising within aggregated
components. The key to achieve this is the projection of these componentwise routing costs
onto the master problem variables to enable their consideration in the network designs
proposed in subsequent iterations. This projection, which is theoretically possible via
Fourier-Motzkin elimination, is practically realized by Benders decomposition. The basis
of the latter framework is to separate over the set of cutting planes describing the projection
in order to avoid the exponential blow-up of adding the complete description to the master
problem.

Aggregation techniques in combination with Benders decomposition have mostly been con-
sidered so far as a means to shrinken the scenario tree for stochastic programming problems.
Publications like Dempster and Thompson (1998), Trukhanov et al. (2010) and Lumbreras
and Ramos (2013) may serve as a reference here. In Pruckner et al. (2014), the authors
develop a model for the optimal expansion of power generation capacities in the German
federal state of Bavaria. The problem is decomposed into a Benders master problem deter-
mining the investments into new capacities, while the subproblem checks whether power
demands can be fullled such that the technical constraints of the power plants are met.
Due to the long-term planning horizon under investigation and the need to represent the
technical constraints on a very ne-grained timescale, the latter is aggregated and only
disaggregated as needed to remedy infeasibilities, which leads to an eective algorithm. In
Shapiro (1984), a Zipkin-like node aggregation scheme is proposed in the context of net-
work design problems. The rules to determine which nodes may be aggregated are quite
restrictive, as they have to full several similarity criteria (cf. Zipkin (1977)). Under these

© Springer Fachmedien Wiesbaden 2016


A. Bärmann, Solving Network Design Problems via Decomposition,
Aggregation and Approximation, DOI 10.1007/978-3-658-13913-1_9
124 8. Integration of Routing Costs into the Aggregation Scheme

assumptions, it is possible to solve the problem via a Benders approach and to establish
error bounds on the obtained solution. Computational results are not presented, however.
The idea for a general purpose Benders algorithm for aggregated formulations can be found
in Gamst and Spoorendonk (2013). The authors propose to set up the Benders master
problem as an aggregated version of the original problem with additional variables mod-
elling the coupling between the coarse and the ne decision space. The original problem
itself is relaxed to a linear program and functions as the Benders subproblem. Interestingly,
the master problem is now solved as an LP, while branching is perfomed on the original
variables to yield an exact algorithm. However, computational results are not available
yet.

In this chapter, we choose an approach dierent to the ones above. We will integrate the
routing costs of the arcs into the aggregation scheme of Chapter 7 by incorporating an
additional subproblem. Its task is to calculate the componentwise routing cost of feasible
solutions and to pass this information to the master problem via a lifted Benders optimality
cut.

In Section 7.1.3, we already saw that the Benders feasibility cuts can be replaced by the
primal cuts of our aggregation scheme, which yields tighter relaxations. Doing so implies
passing from a static subdivision between master problem and subproblem, as it is consid-
ered in a traditional Benders decomposition, to a dynamic subdivision. In the following,
we will derive a hybrid algorithm between aggregation and Benders decomposition that
additionally replaces the original Benders feasibility cuts by suitably lifted versions of them
to account for the shrinking size of the subproblem. Before we go into further details, we
investigate a special case to motivate our algorithmic idea for the integration of routing
costs: the application to minimum-cost s-t-ows. Then we give a full description of our
aggregation algorithm with routing costs as a specialized cutting-plane method together
with the proof of its correctness.

8.1. A Special Case: Aggregating Minimum-Cost s-t-Flow


Problems
To motivate the use of Benders optimality cuts to incorporate the routing costs into our
aggregation scheme, we consider the case of a network ow problem which is solvable in
polynomial time: the minimum-cost s-t-ow problem (see for example Ahuja et al. (1993)).
In this problem, we are to send a given demand d ≥ 0 from a node s to a node t within
a given directed network R
G = (V, A). The arcs are equipped with capacities c : A → +
and routing costs f: A→ R+, and we have to nd a minimum-cost routing respecting the
capacities. In the following, we treat a parameterized version of the problem where the
demand d is not xed in advance and consider the optimal routing cost as a function of
this parameter. It shall be denoted by Φ in the following. The parameterized demand d
can be seen as the ow on a new articial arc from s to t, which could be introduced as
an aggregated substitute for a subnetwork G within a more complex surrounding network.
The value Φ(d) would then describe the routing costs within the subnetwork depending on
the ows in the surrounding network.
8.1 A Special Case: Aggregating Minimum-Cost s-t-Flow Problems 125

Function Φ can be evaluated for a given demand value d ≥ 0 by solving the following linear
program, which describes the minimum-cost s-t-ow problem for network G:
P
Φ(d) = min f a ya
a∈A

P P  d, if v = O(r)
s.t. ya − ya = −d, if v = D(r) (∀v ∈ V )
a∈δv+ a∈δv−

0, otherwise (8.1)

ya ≤ c a (∀a ∈ A)
ya ≥ 0 (∀a ∈ A),

where ya is the ow on arc a∈A and where the constraints are the ow conservation and
the capacity constraints respectively.

Let D≥0 be the maximal throughput of G, i.e. the value of a minimal cut. We would
like to characterize now the optimal value function Φ : {d ∈ R+ | d ≤ D} → R+ of (8.1),
which is the function assigning to each demand value d ≤ D the corresponding optimal
value of Problem (8.1). First, we will give a complete description of Φ via a Benders
reformulation of the minimum-cost ow problem, from which we see that Φ is piecewise
linear. Second, we will explore how to obtain piecewise-linear underestimators for Φ via a
successive-shortest-paths approach.

8.1.1. The Benders Reformulation


We begin by showing that Φ is piecewise linear. We do this by projecting out the original
variables in Problem (8.1), leaving only the demand z. The algorithmic approach to do so
is Benders decomposition.

For a xed amount of ow d to be routed through network G, the dual problem to (8.1) is
given by:
P
max d(αs − αt ) − ca βa
a∈A

s.t. αv − αw − βa ≤ fa (∀a = (v, w) ∈ A) (8.2)

βa ≥ 0 (∀a ∈ A)

where α and β are the dual variables to the ow conservation and the capacity constraints
respectively.

Let E = {αe , β e ) | t = 1, . . . , E} denote the set of extreme points of the feasible set of
Problem (8.2). The restriction of d to values less than or equal to D allows us to disregard
the extremal rays of Problem (8.2). Together, we can write the Benders master problem
as:
min Φ(d)
(8.3)
Φ(d) ≥ (αse − αte )d − βae ca (∀e ∈ E).
P
s.t.
a∈A
126 8. Integration of Routing Costs into the Aggregation Scheme

In a classical Benders algorithm, the way to evaluate Φ(d) is to solve a restricted version of
the master problem (8.3), where E is substituted by a smaller or empty subset, to obtain
a lower bound for Φ(d). This lower bound is used in Problem (8.2) to either prove its
optimality or to yield a cutting plane to rene the master problem. Of interest to us,
however, is the structure of function Φ. According to its representation in Problem (8.3),
it can be written as the maximum over a set of linear functions and is therefore a piecewise-
linear convex function. Each non-dominated inequality among the inequalities

X
Φ(d) ≥ (αse − αte )d − βae ca
a∈A

for eP∈ E denes one of its segments with a slope of (αse − αte ) and an axis intercept at
(0, − a∈A βae ca ). Each Benders iteration gives us one new segment of Φ. The complete
evaluation of Φ(d) possibly requires many iterations to yield all the segments of interest
with respect to a given demand d. We will show next how it is possible to compute the rst
K segments of function Φ for some K ∈ N
via shortest-path evaluations instead. These
can be used to form an underestimator of Φ. In the case of an s-t-subnetwork within a
bigger network, they can be precomputed to speed up the solution of Problem (7.1) via an
aggregation-based algorithmic framework.

8.1.2. Computing Underestimators for Φ


For an ecient algorithm to solve Problem (7.1), it is clearly not desirable to incorpo-
rate the complete representation of Φ as a maximum of linear functions due to the high
number of constraints this requires. Therefore, we will develop an aggregation procedure
augmentented by Benders optimality cuts to solve it in Section 8.2. This approach allows
to incorporate one additional segment of Φ per iteration as needed, which results in a
linear underestimator of Φ. For the special case of Problem (8.1), we can give a combi-
natorial intepretation of this approach by showing how to compute Φ via a combinatorial
algorithm. It is based on a well-known technique for the solution of minimum-cost ow
problems, namely the successive-shortest-paths algorithm.
The method requires the notion of so-called pseudoows. A pseudoow is a variable as-
signment y:A→ R+ which respects the capacity and non-negativity constraints of Prob-
lem (8.1), but which possibly violates its ow conservation constraints. Furthermore, we
need to introduce the residual network G(y) with respect to some pseudoow y . It arises
from the original network G = (V, A) by replacing the arcs a = (u, v) ∈ A by corresponding
+ − +
forward arcs a = (u, v) and backwards arc a = (v, u). Arc a is assigned a routing cost
+
of fa and a residual capacity of ra := ca − ya , while arc a
− bears a routing cost of −f
a

and a residual capacity of ra := ya . The residual network G(x) then contains those arcs
which possess a positive residual capacity.

The main algorithmic idea is then to start with y = 0 as a feasible pseudoow and to
augment it iteratively along shortest paths from s to t in the corresponding residual network
with respect to the routing costs of the arcs. A detailed description of the successive-
shortest-paths algorithm can be found in Ahuja et al. (1993). In the following, we show
that the piecewise-linear representation cost function Φ can be computed by using this
8.1 A Special Case: Aggregating Minimum-Cost s-t-Flow Problems 127

algorithm. More exactly, we show how to compute its rst K segments for some K > 0.
Note that the slopes of these segments increase monotonously, but not necessarily in a
strict fashion as there might be several shortest paths to choose from in each iteration of
the successive-shortest-paths algorithm. Consequently, the order in which the algorithm
computes segments with equal slope depends on the choice of the shortest path in each
iteration.

Theorem 8.1.1. Let δk be the slope of the k -th segment of Φ and let (νk , ωk ) be its support-
ing point, i.e. the rst point on that segment. Then δk is equal to the length of a shortest
path Pk in the residual network in the k -th iteration of the succesive-shortest-paths algo-
rithm while νk is the value of the corresponding pseudoow at the beginning of the iteration.
Furthermore, we have ω1 = 0 and ωk = ωk−1 + Γk · δk for k > 1, where Γk is the minimal
capacity among the arc capacities along Pk .

Proof. For d = 0, an optimal solution to Problem (8.2) is obviously given by (ᾱ, β̄) = (0, 0),
thus ν1 = 0 and ω1 = 0. Let ŷ be a pseudoow in G whose value d ˆ is such that it lies
in the domain of denition of the rst segment of Φ. As this segment contains the origin,
dˆ belongs to a solution to 8.2 with
P
a∈A ca βa = 0. Consequently, ŷ is a solution to
Problem (8.1) where the capacity constraints are non-binding. It follows that ŷ routes all
its ow along a shortest path from s to t, whose cost is (α̂s − α̂t ). The routing cost of ŷ is
ˆ s − α̂t ).
therefore d(α̂

On the other hand, let P1 be a shortest path from s to t and let Γ1 be the minimal
capacity of any of the arcs in P1 . This value Γ1 corresponds to the maximal amount
of ow which can be routed along P1 in any feasible solution to Problem (8.1). Thus,
(Γ1 , Γ1 (α̂s −α̂t )) is the point where the rst segment ends and therefore the supporting point
of the second segment. Together, this shows that the successive-shortest-path algorithm
correctly computes the rst segment of Φ.

Now assume that we have already computed the rst k − 1 segments of Φ, k > 1. In
particular, assume we are given the supporting point of the k -th segment, (νk , ωk ) as well
as a pseudoow ŷ with a value of νk . Let G(ŷ)) be the corresponding residual network to
ŷ , and let ỹ be the increment to ŷ performed in the subsequent iteration of the successive-
shortest-paths algorithm. Then we can substitute y = ŷ + ỹ in Problem (8.1). We obtain:

P
min fa ỹa
a∈A

P P  d − νk , if v = O(r)
s.t. ỹa − ỹa = −d + νk , if v = D(r) (∀v ∈ V )
a∈δv+ a∈δv− 0,

otherwise

ỹa ≤ ca − ŷa (∀a ∈ A)


ỹa ≥ −ŷa (∀a ∈ A),

P
where we leave out the constant a∈A fa ŷa in the objective function. By denition, ỹa is a
free variable. We split it into two bounded variables ỹa+ , ỹa− ≥ 0 and dene ỹa := ỹa+ − ỹa− .
128 8. Integration of Routing Costs into the Aggregation Scheme

This yields:

fa (ỹa+ − ỹa− )
P
min
a∈A

 d − νk , if v = O(r)
(ỹa+ ỹa− ) (ỹa+ ỹa− )
P P
s.t. − − − = −d + νk , if v = D(r) (∀v ∈ V )
a∈δv+ a∈δv− 0,

otherwise

ỹa+ − ỹa− ≤ ca − ŷa (∀a ∈ A)


ỹa+ − ỹa− ≥ −ŷa (∀a ∈ A).

The last two constraints can be simplied because only one of ỹa+ , ỹa− will be non-zero in
an optimal solution and because ŷ ≥ 0 and c − ŷ ≥ 0 holds at the same time, such that
the nal problem reads:

fa (ỹa+ − ỹa− )
P
min
a∈A

 d − νk , if v = O(r)
(ỹa+ ŷa− ) (ŷa+ ŷa− )
P P
s.t. − − − = −d + νk , if v = D(r) (∀v ∈ V )
a∈δv+ a∈δv− 0,

otherwise

ỹa+ ≤ ca − ỹa (∀a ∈ A)


ỹa− ≤ ỹa (∀a ∈ A).
(8.4)

This is exactly the optimal routing problem corresponding to the current residual network
G(ŷ). The variables ỹa+ and ỹa− correspond to the ows on the forward and backward arcs
of the arcs a∈A respectively.

Let y := ŷ + ỹ be a pseudoow whose value νk + d lies in the domain of denition of the


k -th segment of Φ. By denition, ỹ corresponds to a dual solution to Problem (8.4) where
the dual variables to the variable bounds are 0. Thus, the dual problem corresponds to
a shortest-path problem in G(ŷ) and the slope of the k -th segment of Φ is equal to the
length of a shortest path.

On the other hand, we can determine the next supporting point as (νk + Γk , ωk + Γk (α̂s −
α̂t )), Γk is the minimal arc capacity along the shortest path corresponding to the k -
where
th segment of Φ. Together this shows that the successive-shortest-path algorithm correctly
computes the rst k + 1 segments of Φ, which completes the proof.

The above theorem provides a combinatorial view on the Benders optimality cuts used to
underestimate the routing cost function Φ. In the case of s-t-subnetworks, it would even
be possible to use the successive-shortest-path algorithm to precompute the rst segments
of Φ for each aggregated subnetwork to speed up the Benders procedure. However, it is
unclear how to extend this to more complex subnetworks as already the case of multiple
entry or exit nodes makes Φ a multi-dimensional function. In the following, we show how
lifted Benders optimality cuts can be integrated into the aggregation scheme of Chapter 7
to solve general instances of Problem (7.1) with non-zero routing costs.
8.2 An Aggregation Algorithm incorporating Routing Costs 129

8.2. An Aggregation Algorithm incorporating Routing Costs


The aggregation scheme developed in Chapter 7 is an exact method for the solution of
the network design problem (7.1) when there are no routing costs, i.e. f = 0. For non-
zero f, it is straightforward to consider the routing costs of the arcs in the aggregated
graph in the objective function of the master problem. However, this neglects the routing
costs arising within the components of the aggregation, which means a loss of information.
The scheme still delivers feasible solutions in this case, but they need not be optimal any
more. Therefore, we now derive an extension of the aggregation scheme which projects the
routing costs within each component Vi ∈ V ϕ onto the arcs of the master problem graph
Aϕ . The algorithmic means to do so is the introduction of an additional subproblem which
evaluates the costs for a feasible routing within a component and reformulates them as a
cutting plane to be added to the master problem.

According to the discussion in Section 8.1, a suitable adaption of the aggregation master
problem (7.1) for a given clustering function ϕ is as follows:

P P
min k a ua + fa ya + Φ(u, y)
a∈Aϕ a∈Aϕ
P P
s.t. ya − ya = d i (∀i ∈ Vϕ )
a∈δi+ a∈δi−
(8.5)
ya ≤ ca + Ca ua (∀a ∈ Aϕ )
ua ∈ Z+ (∀a ∈ Aϕ )
ya ≥ 0 (∀a ∈ Aϕ ).

This new master problem directly includes the routing costs of the arcs in Aϕ in the
objective function. The routing costs within the components of the aggregated graph
are incorporated via a piecewise-linear cost function Φ as explained above. This func-
tion
P Φ is iteratively underestimated via Benders cutting planes which project the term

a∈A\Aϕ f a ya onto variables ua for a∈A and ya for a ∈ Aϕ .

Let (ū, ȳ) be a solution to the master problem (8.5) which can be extended to a feasible
solution to the original problem (7.1) in the sense of Section 7.1.1, i.e. we are able to nd
a feasible routing within the components. In this case, the routing costs arising within the
components are given as the optimal value of the following minimum-cost ow problem:

P
min f a ya
a∈A\Aϕ

ya = d¯v (∀v ∈ V )
P P
s.t. ya −
a∈δv+ \Aϕ a∈δv− \Aϕ (8.6)

ya ≤ c̄a (∀a ∈ A \ Aϕ )
ya ≥ 0 (∀a ∈ A \ Aϕ ),

d¯v := dv − a∈δv+ ∩Aϕ ȳa + a∈δv− ∩Aϕ ȳa for v ∈ V are the induced
P P
where node demands,
and where c̄a := ca + Ca ūa are the arc capacities within the components.
130 8. Integration of Routing Costs into the Aggregation Scheme

From the dual to Problem (8.6), we can derive the Benders optimality cut for the current
solution (ū, ȳ). It is given by:

P ¯ P
min dv α v − c̄a βa
v∈V a∈A\Aϕ

s.t. αv − αw − βa ≤ fa (∀a = (v, w) ∈ A \ Aϕ ) (8.7)

βa ≥ 0 (∀a ∈ A \ Aϕ )

with dual variables α for the ow conservation constraints and β for the capacity con-
straints. Whenever the current underestimation of the componentwise routing costs Φ(ū, ȳ)
within the master problem is lower than the optimal value of Subproblem (8.7), we can
introduce a Benders optimality cut to update function Φ. By backsubstitution for d¯ and
c̄, this cut can be stated as:

X X X X
ᾱv dv − β̄a ca + (ᾱw − ᾱv )ya − (Ca β̄a )ua ≤ Φ(u, y), (8.8)
v∈V a∈A\Aϕ a=(v,w)∈Aϕ a∈A\Aϕ

given a dual optimal solution (ᾱ, β̄). The interpretation of this cutting plane is that a given
suboptimal solution with respect to the original problem can be improved by rerouting the
ow in the master problem graph or by creating new capacities within the components.

The procedure can start with a coarse, possibly trivial underestimation of Φ. Adding
a violated optimality cut of the form (8.8) to the master problem then yields a better
estimate for the routing costs within the components. The idea is to iterate this until
either no more violated optimality cuts exist or until the induced ow within any of the
components becomes infeasible. In the rst case, we know that we have found an optimal
solution, as the master problem is a relaxation of the original problem. In the latter case, it
would be possible in principle to use Benders feasibility cuts to cut o infeasible solutions.
However, we already now that they are dominated by the cutting planes produced by
the aggregation scheme. Therefore, we perform a disaggregation step which updates the
distribution of the arcs between master problem and subproblem and restart the method.

Altogether, the above considerations lead to a generalization SAGGB of Algorithm SAGG


to the case of non-zero routing costs. We do not expect SAGGB to be very ecient
as computing the correct routing cost entails the addition of many feasibility cuts. A
sequential scheme would start from scratch for each of those. However, it is also possible
to derive analogous generalizations of Algorithms IAGG and HAGG, named IAGGB
and HAGGB respectively, which operate within an integrated branch-and-bound tree.
The important consideration here is to guarantee the global validity of the optimality
cuts. As long as no disaggregation is performed after starting the branch-and-bound phase
of the algorithms, this is no problem as the method then behaves like ordinary Benders
decomposition. After a disaggregation step, however, the cutting plane has to be modied
as the master problem graph grows while the components shrink. The arcs moving from
inside the components to the master graph are no longer valued in the dual Benders
subproblem (8.7). In the extreme case of complete disaggregation, this subproblem is
empty with an optimal value of zero.
8.2 An Aggregation Algorithm incorporating Routing Costs 131

The solution is a suitable lifting of the optimality cuts obtained from Subproblem (8.7).
We add the primal routing cost of all arcs that have passed to the master problem after the
root node to the left side of (8.8). If ϕ is the clustering function corresponding to the initial
aggregation and ϕ̄ is the clustering function describing the current state of aggregation at
a node within the branch-and-bound tree, the new optimality cut reads as follows:
X X X X X
ᾱv dv − β̄a ca + f a ya + (ᾱw − ᾱv )ya − (Ca β̄a )ua ≤ Φ(u, y).
v∈V a∈A\Aϕ̄ a∈Aϕ̄ \Aϕ a=(v,w)∈Aϕ̄ a∈A\Aϕ̄
(8.9)

A solution produced within the branch-and-bound procedure can then be accepted as


feasible if no violated cutting plane of the form (8.9) exists. This can be checked via the
optimality cuts already added to the master problem. Let (ū, ȳ) be the current solution
candidate, let Φ̄ be the estimate of the routing costs within the components according to
the initial clustering ϕ and let Φ(ū, ȳ) be the actual value of the routing costs within the
components according to the current clustering ϕ̄. If
X
Φ̄ − fa ȳa = Φ(ū, ȳ)
a∈Aϕ̄ \Aϕ

holds, the cost estimate is correct, and the solution may be accepted. Otherwise, we add
Cutting Plane (8.9) and reject the solution. Altogether, this leads to a hybrid algorithm
between the original aggregation scheme and Benders decomposition.

All algorithmic considerations of Section 7.2 remain valid except for the global feasibility
test. It has to be dropped as the routing on the master graph is crucial for the induced
routing cost arising within the components and cannot simply be replaced by another
feasible routing. Note that a trivial but powerful heuristic within the branch-and-bound
process is to repair a solution cut o by (8.9) by providing the actual value of Φ for the
routing costs within the components. If its correct cost is still better than that of the best
solution found so far, it can be retained and become the new incumbent.

The following theorem states the correctness of the proposed method. As the general idea
behind the approach has a much wider applicability than the inclusion of the routing costs
into the aggregation scheme, we are able to give a very general proof for the correctness of
a Benders decomposition that allows shifting variables from the subproblem to the master
problem in the process.

For k ≥ 0 and f such that G does not contain negative cycles, SAGGB,
Theorem 8.2.1.
IAGGB and HAGGB are nite and return an optimal solution to the network expansion
problem (7.1) for the original graph G.

Proof. Consider the following optimization problem with variables x ∈ Rn and y ∈ Rm for
n ≥ 0 and m ≥ 1 as well as vectors b and c and matrices A and B of suitable dimensions:
min cT x + dT y
s.t. Ax + By ≤ b
x ≥ 0
y ≥ 0.
132 8. Integration of Routing Costs into the Aggregation Scheme

Now, assume that the y -variables are projected out of the problem. Let x̄ be a choice of
variables x such that a feasible solution (x̄, ȳ) exists for this problem. In this case, the
primal and dual Benders subproblem are

min dT y
s.t. By ≤ b − Ax̄
y ≥ 0

and
max (−b + Ax̄)T α
s.t. −B T α ≤ d
α ≥ 0

respectively, where α are the dual variables to the single constraint. The Benders master
problem simply takes the form min{Φ(x, y) | x ≥ 0}. Note that we write Φ(x, y) for the
estimation of the subproblem cost, i.e. including the y -variables, to allow for the fact that
part of them are reintroduced to the master problem in the course of the solution process.
The Benders optimality cut for an optimal dual solution ᾱ according to x̄ is of the form

Φ(x, y) ≥ −ᾱT b + (AT ᾱ)T x.

We now consider the situation that part of the y -variables are shifted from the subproblem
to the master problem. Without loss of generality, we assume that this is done for the rst
p variables y1 ∈ Rp, while the remaining vector of variables y2 ∈ Rm−p remains in the
subproblem.

If y1 had been part of the master problem from the start, the primal and dual Benders
subproblem would have been

min dT2 y2
s.t. B2 y2 ≤ b − Ax̄ − B1 ȳ1
y2 ≥ 0

and
max (−b + Ax̄ + B1 ȳ1 )T α
s.t. −B2T α ≤ d2
α ≥ 0

respectively, where B and d are split up into B = (B1 , B2 ) and d = (dT1 , dT2 )T according
to the partitioning of y. Here, α are again the dual variables of the single constraint, and
x̄ and ȳ1 are chosen such that there exists a feasible solution (x̄, ȳ1 , ȳ2 ) to the original
problem. The corresponding Benders optimality cut is

Φ(x, y) ≥ −ᾱT b + (AT ᾱ)T x + (B1T ᾱ)T y1 .


8.2 An Aggregation Algorithm incorporating Routing Costs 133

Now, we would like to show that the addition of its lifted version

Φ(x, y) ≥ −ᾱT b + (AT ᾱ)T x + (B1T ᾱ)T y1 + dT1 y1

to the master problem leads to a correct estimation of the subproblem cost after moving
y1 from the subproblem to the master problem. Let Φ̄ be the current estimate for the
subproblem cost according to the initial division of the variables, let ȳ2 be a feasible
completion of the partial solution (x̄, ȳ1 ) and let Φ(x̄, ȳ) be the corresponding objective
value of the shrunk subproblem. If

Φ̄ − dT1 ȳ1 = Φ(x̄, ȳ)

holds, we know that Φ̄ is the exact value for the cost of the initial subproblem incurred by
solution x̄ in the initial master problem. Thus, Φ̄ = dT1 ȳ1 + Φ(x̄, ȳ) is the correct objective
value of the y -variables within the original problem. In the case

Φ̄ − dT1 ȳ1 < Φ(x̄, ȳ),

we have underestimated the cost of solution (x̄, ȳ) and thus add the lifted optimality cut
to correct the estimation.

It remains to show that all the optimality cuts that have been added to the master problem
before reintroducing variables y1 remain valid. In other words, we do never overestimate
the objective value of a solution. This can be seen as follows. The optimality cut that
would have been calculated for x̄ without reintroducing variables y1 to the master problem
dominates all the cutting planes that have been added so far in the point x̄. Thus, the
corresponding cost estimate Φ̂ is at least as high as Φ̄. Let ŷ = (ŷ1 , ŷ2 ) be the optimal
completion calculated by the subproblem according to the initial division of the variables.
Then we have
Φ̄ ≤ Φ̂ = d1 ŷ1 + dT2 ŷ2 ≤ d1 ȳ1 + dT2 ȳ2 .
This leads to
Φ̄ − d1 ȳ1 ≤ dT2 ȳ2 = Φ(x̄, ȳ),
which proves the claim.

It is now easy to see the correctness of the proposed aggregation scheme. It is nite
because disaggregation can only occur until reaching the original graph and because there
is only a nite number of extreme points from which non-dominated optimality cuts can
be derived. The optimality of the returned solution follows from three facts. Firstly,
the non-negativity of k ensures the boundedness of the master problem. Secondly, the
feasible completability of the master problem solution is always ensured by disaggregation
of infeasible components. Thirdly, the absence of negative circles in G with respect to
f ensures that both the primal and dual Benders subproblem are always feasible. This
completes the proof.

The above reasoning does not only show the correctness of the proposed aggregation scheme
with lifted optimality cuts, but also introduces the idea of a generalized Benders decom-
position. In a traditional Benders decomposition, the splitting of the variables between
134 8. Integration of Routing Costs into the Aggregation Scheme

master and subproblem remains xed. Following the outline in the proof of Theorem 8.2,
it is possible to convert Benders decomposition from a pure row generation scheme into a
row-and-column generation scheme. It is an interesting line for future research to assess
the potential of such a method in general. In the following chapter, we give computational
results for the aggregation schemes presented so far. In the light of this discussion, they
may be seen as special cases of the sketched generalized Benders decomposition, where,
in addition, the feasibility cuts are replaced by primal cutting planes from the original
problem.
9. The Computational Impact of
Aggregation

In the previous two chapters, we developed algorithms for the solution of network expansion
problems based on the aggregation of the underlying graph. These include three dierent
realizations of the basic algorithm for the case without routing costs ( SAGG, IAGG,
and HAGG) as well as a generalization of the integrated and the hybrid aggregation
scheme to include routing costs via Benders optimality cuts ( IAGGB and HAGGB). In
the following, we conduct a computational study to assess the eect of graph aggregation
on the performance of the solution process in contrast to solution via a standard solver.
To this end, we have implemented the ve aggregation schemes using the C++-API of
Gurobi and compare them to Gurobi's default branch-and-bound algorithm. The results
on dierent benchmark sets feature signicant savings in computation time by aggregation,
which demonstrates the validity of the approach.

9.1. Computational Setup and Benchmark Instances


We present computational experiments with the aggregation schemes from above on two
dierent sets of benchmark instances. These are random scale-free networks according to
the preferential attachment model (see Albert and Barabási (2002)) and instances derived
from the real-world railway networks investigated in Chapter 5. Results on two more
benchmark sets originating from a DIMACS shortest-path instance (Demetrescu et al.
(2006)) as well as SNDlib instances (Orlowski et al. (2010)) can be found in Bärmann
et al. (2015b).

For the scale-free network instances, we randomly drew the vector d of demands as well as
the vector c of initial arc capacities. The instances from railway network expansion were
adapted to t the single-commodity problem setting by balancing the multi-commodity
demands, retaining the original initial capacities. To demonstrate the extendibility of the
approach to the multi-commodity network expansion problems, we also include preliminary
results for scale-free networks with a few commodities in the case without routing costs.

The initial capacities of each instance were scaled by a constant factor in order to obtain
dierent percentages of initial demand satisfaction l , which was done by solving an auxiliary
network ow problem. The parameter l indicates which portion of the demand can be
routed given the initial state of the network. For dierent instance sizes, varying the initial
capacities has a signicant impact on the solution time and the solvability in general and
is therefore an important parameter for the forthcoming analysis.

© Springer Fachmedien Wiesbaden 2016


A. Bärmann, Solving Network Design Problems via Decomposition,
Aggregation and Approximation, DOI 10.1007/978-3-658-13913-1_10
136 9. The Computational Impact of Aggregation

To assess the inuence of the routing costs on the performance of the aggregation algo-
rithms, we introduce a parameter σ to describe their magnitude. The routing cost fa of
an arc a ∈ A is then given by fa = σka , i.e. it is chosen to be proportional to its expansion
cost. This is motivated by the models for railway network expansion considered in Part I,
where both the expansion cost and the routing cost are proportional to the length of the
corresponding arc. Parameter σ takes values within the set {0, 10−5 , 10−4 , 10−3 , 10−2 }
to which we refer as the cases of no, small, medium, high and very high routing costs
respectively.

Whenever the generation of instances included random elements, we generated ve in-
stances of the same size and demand satisfaction. The solution times then are (geometric)
averages over those ve instances. If only a subset of the ve instances was solvable within
the time limit, the average is taken over this subset only. We also state the number of
instances that could be solved within the time limit.

The results on scale-free networks without routing costs have been obtained on a queuing
cluster of Intel Xeon— E5410 2.33 GHz computers with 12 MB cache and 32 GB RAM.
The employed version of Gurobi was 5.5. All further computational experiments presented
here are more recent and were carried through under a slightly advanced setup, namely
Gurobi 5.6.3 on a queuing cluster of Intel Xeon— E5-2690 v2 3.00 GHz computers with
24 MB cache and 128 GB RAM. Analysis of solution times showed that the advanced
setup leads to an absolute reduction in computation times of about 30 % for the dierent
algorithms, but that their performance relative to each other stays the same.

For IAGG, HAGG, IAGGB and HAGGB, it was necessary to adjust Gurobi's parameter
settings, which involves a more aggressive cutting-plane generation, a focus on improving
the bound and downscaling the frequency of the heuristics. Additionally, as those algo-
rithms use lazy cuts, dual reductions had to be disabled in order to guarantee correctness.
Implementation SAGG uses Gurobi's standard parameter settings. Each job was run on
4 cores and with a time limit of 10 hours. The aggregation schemes are compared to the
solution of the original network expansion integer program by Gurobi's default branch-
and-bound solver with standard parameter settings. These reference solution times are
denoted by MIP. In the case of MIP, experiments with dierent parameter settings did
not lead to considerably better running times.

We begin with the results on the scale-free networks before we present the results on the
railway network instances.

9.2. Computational Results on Scale-Free Networks


The topology of the instances in this benchmark set has been generated according to
a preferential attachment model. It produces so-called scale-free graphs (see Albert and
Barabási (2002)), which are known to represent the evolutionary behaviour of complex real-
world networks well. Starting with a small clique of initial nodes, the model iteratively
adds new nodes. Each new node is connected to m of the already existing nodes. This
parameter m, the so-called neighbourhood parameter, inuences the average node degree.
9.2 Computational Results on Scale-Free Networks 137

We set m=2 in order to generate sparse graphs that resemble infrastructure networks.
We remark that choosing higher values of m did not inuence the results signicantly in
experimental computations. Furthermore, we chose 80 % of the nodes as terminals, i.e.
nodes with non-zero demand, in order to represent a higher but not overly conservative
load scenario. The module capacities for these instances were chosen as 0.25 % of the total
demand in order to obtain reasonable module sizes with respect to the scale of the demand.
Varying these two parameters did not lead to signicantly dierent results either. Finally,
the module costs were drawn randomly.

We begin by comparing SAGG, IAGG and HAGG against MIP in the case of zero routing
costs. In Section 9.2.4, we compare IAGGB and HAGGB against MIP for routing costs
of varying size.

9.2.1. Computational Results for Small Instances


We begin our study by analysing the performance of the aggregation method on small
instances with dierent levels l of initial demand satisfaction. To this end, we generated
random scale-free networks with 100 nodes. Furthermore, we considered a percentage
l ∈ {0, 0.05, 0.1, 0.2, . . . , 0.8, 0.9, 0.95}. In the following, we rst determine which imple-
mentation of the aggregation scheme performs best. In a second step, we compare the best
implementation with MIP.
l IAGG SAGG HAGG

0 3/196.10 3/502.30 3/230.85


0.05 5/22.44 5/186.98 5/63.04
0.1 5/111.15 4/676.27 4/106.75
0.2 5/34.57 5/110.92 5/55.57
0.3 5/8.02 5/25.69 5/10.16
0.4 5/8.01 5/35.03 5/8.47
0.5 5/4.18 5/5.65 5/2.71
0.6 5/5.63 4/4.23 5/6.07
0.7 5/1.17 5/0.81 5/0.82
0.8 5/0.57 5/0.49 5/0.51
0.9 5/0.23 5/0.30 5/0.25
0.95 5/0.12 5/0.21 5/0.18

Table 9.1.: Number of instances solved / average solution times [s] for the three aggregation
algorithms on scale-free networks with 100 nodes and varying values of initial demand
satisfaction l

In Table 9.1, we report the solution times in seconds, each averaged over ve instances
with the same value of l. If not all instances could be solved within the time limit, the
average is taken over the subset of solved instances. The rst number in each cell states the
number of solved instances, whereas the second gives the average solution time in seconds.
138 9. The Computational Impact of Aggregation

Note that we apply the geometric mean for the average values in order to account for
outliers. If a method could not solve any of the ve instances for a given l, we denote
this by an average solution time of  ∞. The fastest method in each row is emphasized in
bold letters. In presence of unsolved instances, we rank the methods rst by the number
of solved instances and second by their average solution time.

The results for the instances from Table 9.1 are also presented as a performance prole
in Figure 9.1. For each aggregation method, it shows the percentage of instances solved

100

90

80

70
% of instances solved

60

50

40

30

20

10 IAGG
SAGG
HAGG
0
1 10 100
Multiple of fastest solution time (log-scale)

Figure 9.1.: Performance prole for the three aggregation algorithms on scale-free networks
with 100 nodes with respect to solution time

within increasing multiples of their shortest solution time. The information deduced from
this kind of diagram is twofold. First, the intercept of each curve with the vertical axis
shows the percentage of instances for which the corresponding method achieves the shortest
solution time. Thus, the method attaining the highest such intercept is the one which
solves the highest number of instances fastest. Second, for each multiple m ≥ 1 on the
horizontal axis, the diagram shows the percentage of instances that a method was able
to solve within m times the shortest solution time achieved by any of the methods. In
particular, the intercept with the vertical axis corresponds to m = 1. Note that in all our
performance proles, the horizontal axis is log-scaled, as it is usually done. The message
of such a diagram is an estimation of the probability with which each method will solve a
given instance fastest, and how good it is in catching up on instances for which it is not the
fastest one. A more detailed introduction to performance proles can be found in Dolan
and Moré (2002).

Altogether, we see that IAGG performs best for the small scale-free networks when com-
pared to SAGG and HAGG. It solves the majority of instances within the shortest solution
9.2 Computational Results on Scale-Free Networks 139

time and solves 97 % of the instances, which is the largest value among the three methods.
Furthermore, there is no instance for which it requires more than four times the shortest
solution time.

In Table 9.2, we thus compare IAGG with MIP.


l MIP IAGG

0 3/1373.11 3/196.10
0.05 5/266.07 5/22.44
0.1 4/904.29 5/111.15
0.2 5/136.43 5/34.57
0.3 5/29.30 5/8.02
0.4 5/28.16 5/8.01
0.5 5/2.40 5/4.18
0.6 5/19.81 5/5.63
0.7 5/0.21 5/1.17
0.8 5/0.08 5/0.57
0.9 5/0.05 5/0.23
0.95 5/0.04 5/0.12

Table 9.2.: Number of instances solved / average solution times [s] of MIP and IAGG for
scale-free networks with 100 nodes and varying values of initial demand satisfaction l

We nd that our aggregation approach is benecial whenever the instance cannot trivially
be solved within a few seconds. For instances with small initial demand satisfaction, we
observe signicantly faster solution times for IAGG, and we see that it is able to solve
more instances within the time limit. Even without any preinstalled capacities (l = 0), the
aggregation scheme attains an average solution time which is six times smaller than that
of MIP. From l = 0.7 upwards, the running times of both algorithms are negligible, and
the tiny advantage for MIP can be attributed to the overhead caused by performing the
aggregation. The superior performance on instances with small initial demand satisfaction
seems surprising. It contrasts the fact that the number of components in the nal state of
network aggregation in IAGG converges to the number of nodes in the original instance.
This is represented in Figure 9.2, where we show the average number of components in the
nal iteration as a function of the percentage of initial demand satisfaction.

Obviously, the aggregation framework performs better than the standard approach MIP
even in case of complete disaggregation. In order to determine what causes this behaviour,
we tested whether the aggregation approach could determine more eective branching de-
cisions within the branch-and-bound procedure. To this end, we increased the branching
priorities in MIP for those variables that enter the master problems of the aggregation
scheme in early iterations. We found that these branching priorities did not lead to bet-
ter running times for MIP. This suggests that the cutting planes generated within the
aggregation procedure are more powerful than the ones generated within MIP.
140 9. The Computational Impact of Aggregation

1
IAGG

0.9

% of nodes of the initial problem 0.8

0.7

0.6

0.5

0.4

0.3

0.2

0.1

0
0 0.05 0.1 0.15 0.2 0.25 0.3 0.35 0.4 0.45 0.5 0.55 0.6 0.65 0.7 0.75 0.8 0.85 0.9 0.95
Initial demand satisfaction (%)

Figure 9.2.: Average number of components for small scale-free networks in the last iteration
of IAGG in relation to the original size of 100 nodes with varying level of initial demand
satisfaction

9.2.2. Behaviour for Medium-Sized and Large Instances

We now consider larger instances within a range of 1000 to 25000 nodes as well as an initial
demand satisfaction from 60 % to 98 %. We applied a selection rule to sort out instances
which are too easy or too hard to solve. We required that for at least three out of
ve instances per class, any of the four considered methods has a solution time between
10 seconds and the time limit of 10 hours. Table 9.3 lists the instances which comply with
this selection rule. Note that the relevant instances can be located mainly on the diagonal
as an increasing instance size requires an increasing level of initial capacities in order to
remain solvable within the time limit.

Figure 9.3 shows the performance prole for the scale-free networks from Table 9.3. In
Figure 9.3a, we see that Implementation HAGG performs best for the instances with up to
5000 nodes and that IAGG is almost as good. Accordingly, these results suggest the choice
of one of those two methods. However, the picture changes for the large instances with at
least 10000 nodes, which have a high level of pre-installed capacities, see Figure 9.3b. Here,
HAGG clearly performs very poorly. Instead, SAGG solves a majority of the instances
fastest (42 %), while IAGG performs only slightly worse. As a result of Figure 9.3, we come
to the conclusion that the overall best choice is IAGG, as it is not much worse than HAGG
on the medium-sized instances and much better on the large networks. Furthermore, it
outperforms both other implementations when considering small multiples of the shortest
solution times.
9.2 Computational Results on Scale-Free Networks 141

l
|V | 0.6 0.65 0.7 0.75 0.8 0.85 0.9 0.92 0.94 0.96 0.98
1000 X X X X X X
2000 X X X X X
3000 X X X X X
4000 X X X X X
5000 X X X X X X
10000 X X X X X
15000 X X
20000 X
25000 X

Table 9.3.: Instance sizes |V | and initial demand satisfactions l that comply with the
selection rule for medium-sized and large instances are marked by an X in the table.

100 100

90 90

80 80

70 70
% of instances solved

% of instances solved

60 60

50 50

40 40

30 30

20 20

10 IAGG 10 IAGG
SAGG SAGG
HAGG HAGG
0 0
1 10 100 1 10 100
Multiple of fastest solution time (log-scale) Multiple of fastest solution time (log-scale)

(a) Medium-sized networks with up to 5000 nodes (b) Large networks from 10000 nodes on

Figure 9.3.: Performance proles for the medium-sized and large scale-free networks from
Table 9.3, comparing the three aggregation schemes

The comparison between IAGG and MIP on the instances from Table 9.3 is shown in
Table 9.4. The instances are grouped by initial demand satisfaction. We see that IAGG is
better comparing the average solution times for almost all instances under consideration.
A special remark is to be made on the total number of solved instances, which is the rst
number in each cell. Here, we see that within the time limit,IAGG can always solve at
least as many instances as MIP, often more. Furthermore, we note that even though the
number of solved instances is larger for IAGG, the geometric mean of the solution times
is still lower compared to the solution times of MIP. Thus, IAGG solves the instances
signicantly faster than the standard MIP approach. These statements are underlined
by the performance prole over the same instances, which is shown in Figure 9.4. The
intgerated aggregation scheme IAGG clearly outperforms the standard approach MIP as
it solves about 86 % of all instances fastest.
142 9. The Computational Impact of Aggregation

|V | l MIP IAGG
|V | l MIP IAGG
1000 0.75 4/142.31 5/92.40
4000 0.92 5/979.59 5/51.89
2000 0.75 3/6715.29 4/1583.72
5000 0.92 5/1499.94 5/61.39
1000 0.8 4/52.75 5/34.25 10000 0.92 0/∞ 3/6139.25
2000 0.8 5/944.84 5/143.04
3000 0.94 5/21.08 5/13.17
3000 0.8 2/6605.03 3/1806.96
4000 0.94 5/121.86 5/21.30
1000 0.85 5/19.68 5/9.31 5000 0.94 5/450.84 5/36.23
2000 0.85 5/140.75 5/43.15 10000 0.94 3/9004.99 5/338.51
3000 0.85 5/2783.47 5/397.87
3000 0.96 5/5.62 5/11.76
4000 0.85 3/14164.16 5/884.52
4000 0.96 5/30.08 5/8.76
5000 0.85 1/31214.39 3/5484.70
5000 0.96 5/65.41 5/20.21
2000 0.9 5/24.39 5/12.78 10000 0.96 4/4012.20 5/109.48
3000 0.9 5/340.59 5/35.94 15000 0.96 2/25343.05 5/2395.53
4000 0.9 5/1433.56 5/115.57
5000 0.98 5/6.70 5/9.53
5000 0.9 5/3113.36 5/195.83
10000 0.98 5/421.14 5/33.07
10000 0.9 0/∞ 3/11951.60
15000 0.98 5/3690.52 5/90.50
2000 0.92 5/11.32 5/10.14 20000 0.98 5/10982.81 5/470.47
3000 0.92 5/63.97 5/24.38 25000 0.98 2/22687.73 4/4372.76

Table 9.4.: Number of instances solved / average solution times [s] of MIP and IAGG for
the scale-free networks from Table 9.3 grouped by number of nodes |V | and initial demand
satisfaction l

100

90

80

70
% of instances solved

60

50

40

30

20

10
MIP
IAGG
0
1 10 100
Multiple of fastest solution time (log-scale)

Figure 9.4.: Performance prole comparing MIP and IAGG on all medium-sized and large
instances from Table 9.3
9.2 Computational Results on Scale-Free Networks 143

To investigate why the aggregation scheme solves the instances so much faster, we examine
the average number of network components in the nal iteration for four selected instance
sizes, namely |V | ∈ {1000, 2000, 3000, 4000}, as this number strongly inuences the size
of the aggregated network design problem; see Figure 9.5. The results are comparable to

0.9
1000
2000
3000
0.8 4000

0.7
% of nodes of the initial problem

0.6

0.5

0.4

0.3

0.2

0.1

0
0.6 0.65 0.7 0.75 0.8 0.85 0.9 0.92 0.94 0.96
Initial demand satisfaction (%)

Figure 9.5.: Average number of components in the last iteration of IAGG in relation to
the original number of nodes |V | ∈ {1000, 2000, 3000, 4000} for scale-free networks with
varying level of initial demand satisfaction

those for the scale-free networks with 100 nodes as shown in Figure 9.2. Due to the larger
size, these instances could only be solved for higher levels of initial demand satisfaction;
for example, they had to be at least 60 % for graphs with 1000 nodes. The plot shows
that the aggregation algorithms can indeed reduce the number of nodes signicantly when
compared to the number of nodes in the original graph. Consequently, the NDP master
problem remains much smaller than the original NDP. It is also interesting to note that for
the tested instances, the factor by which the number of components is smaller compared to
the original graph increases with the number of nodes. That means that for the large-scale
instances, the number of components is signicantly smaller than the original number of
nodes. Therefore, IAGG obtains a big advantage in computation times and achieves a
higher number of solved instances as we saw [Link] an example, Table 9.5 presents the
results for the instances with V = 3000 nodes. For dierent values of l, average solution
times of the aggregation schemes are compared with those of MIP.
In total, these results for medium and large instances conrm our ndings for small in-
100 nodes. Namely, MIP is vastly outperformed by the aggregation schemes.
stances with
IAGG generally performs best with respect to the number of solved instances and with
respect to the solution times.
144 9. The Computational Impact of Aggregation

l MIP IAGG SAGG HAGG

0.8 2/6605.03 3/1806.96 2/5914.47 2/1299.07


0.85 5/2783.47 5/397.87 5/3761.89 5/623.59
0.9 5/340.59 5/35.94 5/58.96 5/28.68
0.92 5/63.97 5/24.38 5/16.23 5/18.00
0.94 5/21.08 5/13.17 5/9.84 5/10.44
0.96 5/5.62 5/11.76 5/7.99 5/7.81

Table 9.5.: Number of instances solved / average solution times [s] of MIP and the aggre-
gation algorithms for scale-free networks with 3000 nodes and initial demand satisfaction l

9.2.3. Computational Results on Scale-Free Networks with


Multi-Commodity Demand
To demonstrate the wider applicability of the proposed aggregation framework, we show
that Methods SAGG, IAGG and HAGG can be extended to the case of multiple com-
modities in a straightforward fashion. The only thing that changes is that the subproblem
now becomes a multi-commodity network ow problem for each component. We present
results on small scale-free networks with |V | = 100 and l = 0.8 as well as medium-sized
scale-free networks with |V | = 3000 and l = 0.96. The short solution times for these net-
work sizes obtained in the single-commodity case allow for a multi-commodity study with
a varying number of commodities b, for which the demands were again randomly drawn.
Table 9.6 compares the number of instances solved and the average solution times obtained
by MIP and the three implementations of our aggregation scheme for the small instances
with up to 25 commodities. We see that an increasing number of commodities increases

b MIP IAGG SAGG HAGG

5 5/7.13 5/9.72 5/6.78 5/6.93


10 5/258.43 5/52.47 5/126.54 5/86.12
15 4/166.98 5/147.35 5/334.15 5/140.81
20 3/966.64 4/337.68 4/1779.86 4/363.40
25 3/2358.80 5/876.46 3/2678.55 4/2258.37

Table 9.6.: Number of instances solved / average solution time [s] of MIP and the aggrega-
tion algorithms for scale-free multi-commodity networks with 100 nodes, an initial demand
satisfaction of 80 % and an increasing number of commodities b

the diculty of the problem signicantly. On the one hand, this is due to the higher
computational complexity of the multi-commodity subproblems. On the other hand, we
also observe that the graphs tend to be disaggregated much further when optimality can
be proved. Nevertheless, Implementations IAGG and HAGG both outperform MIP, and
for a higher number of commodities, IAGG is preferable. The performance prole for the
9.2 Computational Results on Scale-Free Networks 145

latter in comparison to MIP is shown in Figure 9.6. It underlines these ndings as 70 %


of the instances are solved fasted by IAGG.
100

90

80

70
% of instances solved

60

50

40

30

20

10
MIP
IAGG
0
1 10 100
Multiple of fastest solution time (log-scale)

Figure 9.6.: Performance prole comparing MIP and IAGG on scale-free multi-commodity
networks with 100 nodes

The corresponding results for the medium-sized instances in Table 9.7 and Figure 9.7
show a similar picture. The dimension of the graph allows for fewer commodities to be
considered, but the instances are still best solved by the aggregation schemes. In this case,
IAGG does not only solve more instances to optimality than MIP; there even is no single
instance which is solved faster by MIP than by IAGG.

b MIP IAGG SAGG HAGG

2 5/120.01 5/54.13 5/59.98 5/50.08


3 5/337.98 5/152.21 5/156.87 5/132.78
5 5/4436.23 5/620.31 4/863.54 5/533.49
7 3/15302.56 4/2249.79 1/15708.12 5/5169.41
10 0/∞ 2/5465.29 0/∞ 0/∞

Table 9.7.: Number of instances solved / average solution time [s] of MIP and the ag-
gregation algorithms for scale-free multi-commodity networks with 3000 nodes, an initial
demand satisfaction of 96 % and an increasing number of commodities b

We conclude that our aggregation approach can also be used successfully in the multi-
commodity case. We restricted ourselves to a relatively small number of commodities. For
146 9. The Computational Impact of Aggregation

a larger number of commodities, further algorithmic enhancements should be included, such


as methods that aggregate commodities, in addition to aggregating the network topology.

100

90

80

70
% of instances solved

60

50

40

30

20

10
MIP
IAGG
0
1 10 100
Multiple of fastest solution time (log-scale)

Figure 9.7.: Performance prole comparing MIP and IAGG on scale-free multi-commodity
networks with 3000 nodes

9.2.4. Computational Results on Scale-Free Networks with Routing Costs

Our last experiment on the scale-free networks consists of computations for instances with
100 nodes in the case with routing costs. They include a comparison of the two aggrega-
tion algorithms IAGGB and HAGGB of which we will compare the better one with the
reference algorithm MIP, i.e. solution via Gurobi's standard branch-and-bound algorithm.
This is done for varying initial demand coverage between 0 % and 95 % as in Section 9.2.1
and varying routing costs as described in Section 9.1.

We begin with the comparison of IAGGB and HAGGB. Figure 9.8 shows their behaviour
for small, medium, high and very high routing costs in four performance diagrams. We
see that no clear winner can be determined in the rst three cases. The performance of
Methods IAGGB and HAGGB is very similar, although IAGGB is able to solve slightly
more instances within the time limit. In the last case of very high routing costs, however,
it becomes visible that HAGGB signicantly outperforms IAGGB. The most probable
explanation for this nding is that HAGGB starts with a higher degree of disaggregation
of the underlying graph, which means that more arcs have their routing costs represented
in the master problem. This decreases the number of necessary Benders optimality cuts in
the course of the solution process, which seems benecial.
9.2 Computational Results on Scale-Free Networks 147

100 100

90 90

80 80

70 70
% of instances solved

% of instances solved
60 60

50 50

40 40

30 30

20 20

10 10
IAGGB IAGGB
HAGGB HAGGB
0 0
1 10 100 1 10 100
Multiple of fastest solution time (log-scale) Multiple of fastest solution time (log-scale)

(a) Small routing costs (b) Medium routing costs


100 100

90 90

80 80

70 70
% of instances solved

% of instances solved
60 60

50 50

40 40

30 30

20 20

10 10
IAGGB IAGGB
HAGGB HAGGB
0 0
1 10 100 1 10 100
Multiple of fastest solution time (log-scale) Multiple of fastest solution time (log-scale)

(c) High routing costs (d) Very high routing costs

Figure 9.8.: Performance proles comparing Methods IAGGB and HAGGB on scale-free
networks with 100 nodes for dierent sizes of the routing costs

As HAGGB emerges as the more stable method for varying size of the routing costs, we
MIP. The results, again for all
compare it in the following to the reference implementation
four sizes of the routing costs are given in Table 9.8. The picture is very similar to that in
Table 9.2, which compares IAGG to MIP in the case without routing costs. For smaller
values of l , HAGGB clearly outperforms MIP for any size of the routing costs. For higher
values of l , the solution times become very small, such that the overhead of HAGGB does
not pay o any more. In Table 9.8d for very high routing costs, we recognize a tendency
that MIP catches up and that the solution times of HAGGB are signicantly higher for
high values of l. This trend can be expected to continue for even higher routing costs,
as there comes a point where the objective function is vastly dominated by the routing
costs and the expansion costs are low in comparison. In the extreme case, the problem can
practically be regarded as a continuous minimum-cost ow problem. This type of problem
is easy to solve for MIP, while it requires many Benders cutting planes in HAGGB.
Altogether, we could successfully extend the basic aggregation scheme to the case with
routing costs for the small scale-free networks. These results are complemented by the
corresponding results on railway networks in Section 9.3.2, which possess somewhat larger
underlying graphs.
148 9. The Computational Impact of Aggregation

l MIP HAGGB l MIP HAGGB

0 3/438.97 3/127.37 0 3/595.20 3/165.56


0.05 5/100.40 5/15.23 0.05 5/76.66 5/11.35
0.1 4/277.60 5/133.60 0.1 4/151.25 5/60.68
0.2 5/87.41 5/32.70 0.2 5/50.85 5/19.32
0.3 5/33.91 5/9.61 0.3 5/16.22 5/11.94
0.4 5/29.80 5/5.59 0.4 5/15.53 5/5.70
0.5 5/2.11 5/5.31 0.5 5/1.89 5/1.78
0.6 4/2.52 4/1.86 0.6 5/6.75 4/1.26
0.7 5/0.13 5/0.38 0.7 5/0.13 5/0.42
0.8 5/0.05 5/0.23 0.8 5/0.05 5/0.22
0.9 5/0.03 5/0.10 0.9 5/0.02 5/0.10
0.95 5/0.02 5/0.07 0.95 5/0.02 5/0.07
(a) Small routing costs (b) Medium routing costs

l MIP HAGGB l MIP HAGGB

0 3/111.43 3/60.84 0 3/20.28 5/104.46


0.05 5/72.49 5/8.59 0.05 5/4.51 5/1.98
0.1 5/212.93 5/32.33 0.1 5/6.72 5/2.18
0.2 5/95.67 5/21.20 0.2 5/6.96 5/1.73
0.3 5/27.96 5/9.43 0.3 5/5.11 5/3.08
0.4 5/16.66 5/5.63 0.4 5/4.21 5/1.98
0.5 5/3.30 5/1.73 0.5 5/2.30 5/1.69
0.6 5/1.53 5/1.43 0.6 5/0.95 5/2.39
0.7 5/0.13 5/0.73 0.7 5/0.14 5/2.92
0.8 5/0.06 5/0.75 0.8 5/0.08 5/4.35
0.9 5/0.03 5/0.30 0.9 5/0.06 5/11.59
0.95 5/0.02 5/0.17 0.95 5/0.03 5/8.96
(c) High routing costs (d) Very high routing costs

Table 9.8.: Number of instances solved / average solution times [s] of MIP and HAGGB
on scale-free networks with 100 nodes and varying values of l for dierent sizes of the
routing costs

9.3. Computational Results on Real-World Railway Networks

Our second type of benchmark instances is derived from the real-world railway network
topologies considered in Part I. They comprise the complete German railway network
as well as subinstances of it. These instances vary in size from 62 nodes for the smallest
network up to 1620 nodes for the complete network as summarized in Table 5.2 on page 88.
We used the original arc and module capacities and the demand pattern of the target year
9.3 Computational Results on Real-World Railway Networks 149

2030, which was adapted to single-commodity problem data by considering the net demand
of each node. As a result, these instances do not contain any random elements.

The instances are ltered according to criteria too easy and too hard as in Section 9.2.2.
The upper bound is again given by the time limit of 10 h, while the lower bound is increased
from 10 s to 100 s. This is done in both cases excluding and including routing costs to
enable a unied presentation and to account for the fact that the aggregation schemes
entail a larger overhead in the latter case and start to pay o on somewhat more dicult
instances.

9.3.1. Railway Instances without Routing Costs

We begin our experiments with a comparison of IAGG, SAGG, and HAGG for the case
without routing costs. The results are presented as a performance diagram in Figure 9.9.

100 100

90 90

80 80

70 70
% of instances solved

% of instances solved

60 60

50 50

40 40

30 30

20 20

10 IAGG 10
SAGG MIP
HAGG IAGG
0 0
1 10 100 1 10 100
Multiple of fastest solution time (log-scale) Multiple of fastest solution time (log-scale)

Figure 9.9.: Performance prole compar- Figure 9.10.: Performance prole com-
ing the three basic aggregation schemes paring MIP and IAGG on railway in-
on railway instances stances

We see that Methods IAGG and HAGG both solve the same percentage of instances
fastest. However, IAGG is able to solve all instances under consideration within the given
time limit while HAGG still solves about 90 % of them. Implementation SAGG is clearly
not competitive.

The overall best of the three algorithms, IAGG, is now compared to MIP on the same set
of instances. Figure 9.10 shows the corresponding performance diagram. IAGG clearly
outperforms MIP as it solves more than 90 % of the instances fastest. Furthermore, it is
MIP only solves somewhat less than 60 % of them. A
able to solve all the instances, while
detailed summary of the computation times is given in Table 9.9. It conrms that IAGG
solves almost all the instances faster than MIP, sometimes signicantly faster, while it
is not decisively slower on the few others. Thus, we see that aggregation is at least as
favourable on the real-world networks as on the random networks from before.
150 9. The Computational Impact of Aggregation

Instance l MIP IAGG Instance l MIP IAGG


B-BB-MV 0.5 1353.06 199.33 NRW 0.7 140.04 24.31
B-BB-MV 0.55 209.06 218.97 NRW 0.75 224.81 18.79
BA-BW 0.3 ∞ 8234.56 NS-BR 0.5 12003.63 3035.58
BA-BW 0.4 1231.68 19.63 NS-BR 0.55 ∞ 24310.02
BA-BW 0.45 ∞ 13.89 NS-BR 0.6 1404.99 1826.26
BA-BW 0.5 15.06 8.48 NS-BR 0.65 75.39 24.73
Deutschland 0.7 ∞ 6898.27 TH-S-SA 0.4 ∞ 29243.07
Deutschland 0.75 ∞ 65.08 TH-S-SA 0.45 ∞ 4176.81
HE-RP-SL 0.4 ∞ 5742.33 TH-S-SA 0.5 ∞ 210.66
HE-RP-SL 0.45 ∞ 2786.14 TH-S-SA 0.55 ∞ 1167.42
HE-RP-SL 0.5 2336.17 73.03 TH-S-SA 0.6 1842.68 81.69
HE-RP-SL 0.55 2909.57 33.08 TH-S-SA 0.65 133.27 52.72
NRW 0.6 ∞ 8378.55 TH-S-SA 0.7 40.28 20.12
NRW 0.65 5730.42 497.96 TH-S-SA 0.75 39.30 13.06

Table 9.9.: Solution times [s] for MIP and IAGG on the railway instances with initial
demand satisfaction of l percent

9.3.2. Railway Instances including Routing Costs


Our computational experiments are concluded by assessing the behaviour of the generalized
aggregation schemes IAGGB and HAGGB on the railway network instances including
routing costs.

Figure 9.11 shows their performance diagrams under varying inuence of the routing cost
parameter. The picture is somewhat clearer as for the scale-free networks as HAGGB
performs consistently better from medium-scale routing costs on. For small routing costs,
IAGGB might be preferred as it solves more instances overall, which shows that this case
is closely related to the case featuring no routing costs at all. Note that about 7 % of
these computations were not nished due to numerical diculties within Gurobi which
might have been induced by the Benders cutting planes. When we compare the victorious
aggregation scheme HAGGB to MIP in the following, the corresponding instances are
treated as unsolved for HAGGB.
100 100

90 90

80 80

70 70
% of instances solved

% of instances solved

60 60

50 50

40 40

30 30

20 20

10 10
IAGGB IAGGB
HAGGB HAGGB
0 0
1 10 100 1 10 100
Multiple of fastest solution time (log-scale) Multiple of fastest solution time (log-scale)

(a) Small routing costs (b) Medium routing costs


9.3 Computational Results on Real-World Railway Networks 151

100 100

90 90

80 80

70 70
% of instances solved

% of instances solved
60 60

50 50

40 40

30 30

20 20

10 10
IAGGB IAGGB
HAGGB HAGGB
0 0
1 10 100 1 10 100
Multiple of fastest solution time (log-scale) Multiple of fastest solution time (log-scale)

(c) High routing costs (d) Very high routing costs

Figure 9.11.: Performance proles comparing Methods IAGGB and HAGGB on the rail-
way instances for dierent sizes of the routing costs

Figure 9.12 shows four performance diagrams comparing HAGGB and MIP for the dif-
ferent choices of the routing cost parameter.

100 100

90 90

80 80

70 70
% of instances solved

% of instances solved

60 60

50 50

40 40

30 30

20 20

10 10
MIP MIP
HAGGB HAGGB
0 0
1 10 100 1 10 100
Multiple of fastest solution time (log-scale) Multiple of fastest solution time (log-scale)

(a) Small routing costs (b) Medium routing costs


100 100

90 90

80 80

70 70
% of instances solved

% of instances solved

60 60

50 50

40 40

30 30

20 20

10 10
MIP MIP
HAGGB HAGGB
0 0
1 10 100 1 10 100
Multiple of fastest solution time (log-scale) Multiple of fastest solution time (log-scale)

(c) High routing costs (d) Very high routing costs

Figure 9.12.: Performance proles comparing Methods MIP and HAGGB on the railway
instances for dierent sizes of the routing costs
152 9. The Computational Impact of Aggregation

We see that HAGGB is able to solve a majority of the instances fastest in all the cases
with clear advantages for small, medium, and high routing costs. We also recognize the
same tendency as for the scale-free networks with 100 nodes which lets MIP catch up
for higher routing costs, although it is somewhat concealed by the numerical problems
aecting HAGG especially in Figure 9.12a. The decreasing diculty of the instances for
MIP
higher routing costs can be seen from the increasing percentage of instances solved by
MIP increases the share of instances that it solves fastest.
within the time limit, and

A complete summary of the computation times of MIP and HAGGB is given in Table 9.10,
where instances causing numerical diculties are marked by  −. Note that the varying
number of instances for each magnitude of the routing costs is due to our lter criterion for
overly easy or hard instances. It supports the observations from the performance diagrams
by underlining the strong performance of HAGGB on most instances. The aggregation
scheme is largely favoured up to high routing costs and still competitive for very high
routing costs.

Altogether, we have shown that aggregation is a powerful means to reduce the solution
times of network design problems without sacricing the optimality of the obtained result.
Moreover, the basic idea developed in Chapter 7 can be extended to dierent problem
setting occuring in real-world problem instances.
9.3 Computational Results on Real-World Railway Networks 153

Instance l MIP HAGGB


Instance l MIP HAGGB
B-BB-MV 0.5 731.89 507.45
B-BB-MV 0.55 198.20 128.31 B-BB-MV 0.5 1368.97 768.31
B-BB-MV 0.65 3.71 10.08 B-BB-MV 0.55 170.00 189.91
BA-BW 0.35 ∞ ∞ B-BB-MV 0.65 3.48 8.68
BA-BW 0.4 742.08 14.83 BA-BW 0.4 4017.62 18.50
BA-BW 0.45 ∞ 51.96 BA-BW 0.45 ∞ 109.89
BA-BW 0.55 143.55 7.46 BA-BW 0.5 16.69 5.20
Deutschland 0.75 ∞ 330.56 Deutschland 0.7 ∞ 29928.14
HE-RP-SL 0.4 ∞ ∞ Deutschland 0.75 ∞ 252.88
HE-RP-SL 0.45 ∞ 5107.47 Deutschland 0.8 12.83 104.11
HE-RP-SL 0.5 566.04 33.23 HE-RP-SL 0.4 ∞ 722.55
HE-RP-SL 0.55 775.01 324.97 HE-RP-SL 0.45 ∞ 2723.46
HE-RP-SL 0.6 5.75 − HE-RP-SL 0.5 802.73 20.17
NRW 0.65 12329.86 2471.76 HE-RP-SL 0.55 2729.76 44.23
NRW 0.7 169.42 − NRW 0.65 2417.67 1322.88
NRW 0.75 139.36 41.54 NRW 0.7 345.11 49.77
NS-BR 0.5 5441.22 436.49 NRW 0.75 120.20 30.49
NS-BR 0.55 ∞ 5396.39 NS-BR 0.5 ∞ 3715.65
NS-BR 0.6 2709.92 440.91 NS-BR 0.55 4701.26 4137.90
NS-BR 0.65 16.18 − NS-BR 0.6 3117.81 327.65
NS-BR 0.7 1.06 − NS-BR 0.65 121.52 22.70
TH-S-SA 0.4 31210.83 − TH-S-SA 0.4 ∞ 9072.72
TH-S-SA 0.45 ∞ 21752.85 TH-S-SA 0.45 ∞ 18781.02
TH-S-SA 0.5 4650.49 701.91 TH-S-SA 0.5 2929.61 2321.07
TH-S-SA 0.55 ∞ − TH-S-SA 0.55 ∞ ∞
TH-S-SA 0.6 1712.90 1142.46 TH-S-SA 0.6 2220.04 627.63
TH-S-SA 0.65 118.79 130.10 TH-S-SA 0.65 176.76 118.90
TH-S-SA 0.8 23.13 23.07 TH-S-SA 0.8 17.52 17.42

(a) Small routing costs (b) Medium routing costs

Instance l MIP HAGGB

B-BB-MV 0.45 7428.37 449.29


B-BB-MV 0.5 319.19 117.30
BA-BW 0.4 101.13 15.57
Instance l MIP HAGGB BA-BW 0.45 4313.88 368.54
BA-BW 0.5 ∞ 24318.80
B-BB-MV 0.5 1565.10 521.98 BA-BW 0.55 87.34 −
B-BB-MV 0.55 84.64 120.58 BA-BW 0.6 79.90 2263.22
B-BB-MV 0.65 2.44 8.27 BA-BW 0.7 116.79 29.06
BA-BW 0.3 ∞ 14012.36 BA-BW 0.75 0.59 13.45
BA-BW 0.4 310.06 13.57 Deutschland 0.75 3248.31 −
BA-BW 0.55 246.18 7.69 Deutschland 0.8 92.13 −
BA-BW 0.75 0.51 6.79 Deutschland 0.85 24.95 −
Deutschland 0.75 ∞ 1323.57 Deutschland 0.9 1.48 −
Deutschland 0.8 9.15 − HE-RP-SL 0.4 ∞ 161.92
Deutschland 0.85 4.50 − HE-RP-SL 0.45 3493.37 231.90
HE-RP-SL 0.4 ∞ 791.89 HE-RP-SL 0.5 1161.44 53.22
HE-RP-SL 0.45 ∞ 5197.57 HE-RP-SL 0.55 78.38 24.58
HE-RP-SL 0.5 101.34 25.20 HE-RP-SL 0.65 1.18 9.43
HE-RP-SL 0.55 1911.69 38.96 HE-RP-SL 0.75 0.82 5.13
HE-RP-SL 0.75 0.63 − NRW 0.5 ∞ 1565.48
NRW 0.5 ∞ 16212.22 NRW 0.6 1599.58 1068.06
NRW 0.6 ∞ 28241.39 NRW 0.65 565.79 2165.09
NRW 0.65 2019.42 4258.36 NRW 0.7 56.02 194.42
NRW 0.7 1000.37 193.74 NRW 0.9 0.21 −
NRW 0.75 105.39 11.95 NS-BR 0.5 1803.06 528.92
NS-BR 0.5 8044.32 375.19 NS-BR 0.55 493.40 32.17
NS-BR 0.55 ∞ 4925.51 NS-BR 0.6 257.26 29.58
NS-BR 0.6 917.58 212.88 NS-BR 0.85 0.13 2.85
TH-S-SA 0.4 ∞ 7607.41 TH-S-SA 0.4 ∞ 2861.21
TH-S-SA 0.45 8199.26 11189.70 TH-S-SA 0.45 251.13 206.33
TH-S-SA 0.5 ∞ 426.97 TH-S-SA 0.5 3796.55 132.21
TH-S-SA 0.55 6519.13 ∞ TH-S-SA 0.55 4524.58 64.43
TH-S-SA 0.6 772.52 257.78 TH-S-SA 0.6 628.44 79.17

(c) High routing costs (d) Very high routing costs

Table 9.10.: Solution times [s] of MIP and HAGGB on railway instances with varying
values of l for dierent sizes of the routing costs
Part III.

Approximate Second-Order Cone


Robust Optimization
III. Approximate Second-Order Cone Robust Optimization 157

In practical applications, input data to optimization problems often results from mea-
surements, estimations or forecasts, which are all prone to errors. These errors can have
signicant impact on the feasibility and the quality of the solutions obtained from the
underlying model. This is especially true in the context of infrastructure development,
where demand forecasts are the foundation to plan extensive investments to accommodate
future trac ows. As these forecasts fundamentally shape the design of the future net-
work, incorrect prediction may lead to an inecient routing of the trac and congestions
in one part of the network and unused overcapacities in other regions. Furthermore, such
forecasts typically become less reliable the longer the planning horizon is.

As a remedy for this problem, dierent concepts of optimization under uncertainty have
been devised. A frequently chosen approach in the case of network design problems is
robust optimization under ellipsoidal uncertainty. Over stochastic optimization, it has the
advantage of a much better computational tractability. At the same time, ellipsoidal uncer-
tainty sets allow for the incorporation of statistical data based on past trac observations
as well as probabilistic assumptions on its future development, especially correlation data.
This is an important advantage over polyhedral uncertainty concepts, such as Γ-robustness,
which lack this possibility. However, ellipsoidal uncertainty sets come with the disadvan-
tage that the corresponding robust counterpart turns into a conic-quadratic optimization
problem, which is especially costly in the presence of integer decision variables.

This is where our idea of approximate second-order cone robust optimization comes into
play. We devise a modelling framework that uses a compact and tight approximation
of ellipsoidal uncertainty sets by polyhedra. This way, the arising approximate robust
counterpart stays a linear program. This approach allows to combine the computational
advantages of polyhedral uncertainty sets with the modelling power of ellipsoidal uncer-
tainty sets. Our results show that the so-obtained approximate robust counterpart allows
to compute almost optimal solutions in much shorter time than needed for solving the
exact quadratic model in many cases. Moreover, a coarser approximation can be used to
obtain high-quality solutions quickly.

Part III is structured as follows. In Chapter 10, we describe the basic concept of our
approximation method and give a classication within the context of the existing literature.
Chapter 11 begins with a review of the outer approximation of the second-order cone by
Ben-Tal and Nemirovski. Based on their work, we develop an inner approximation of
the second-order cone which features the same approximation guarantees as the known
outer approximation and which is optimal with respect to the number of variables and
constraints. Then we show how outer and inner approximation can be use to obtain upper
bounds for second-order cone minimization programs. This allows us to derive possible
formulations of an approximate robust counterpart under ellipsoidal uncertainty. The
computational eciency of our approach is demonstrated in Chapter 12 on instances from
portfolio optimization and railway network expansion.
10. Motivation

Robust optimization is an often-used method to deal with uncertainty in the input data to
practical optimization problems. In the case of long-term infrastructure network design,
it is commonly used to obtain a nal network conguration which is suitable for dierent
future trac scenarios. Predictions of trac patterns are typically based on past obser-
vations from which probabilistic assumptions on the future development are derived via
statistic methods. The estimated demand gures may often be described using ellipsoidal
probability distributions, e.g. normal distributions. In the context of robust optimization,
these estimations naturally lead to ellipsoidal uncertainty sets to describe the possible
future trac scenarios.

Ellipsoidal uncertainty sets for uncertain linear programs correspond to a robust coun-
terpart which takes the form of a second-order cone program (SOCP). There are mainly
two approaches used in state-of-the-art solvers to deal with this type of problem, namely
gradient-style approximations of the second-order cone and interior-point (IP) methods.
Both entail signicant computational disadvantages if integer variables are present in the
problem. As interior-point methods lack the possibility of using warm-start for the solution
of the nodes in the branch-and-bound tree, they are uncompetitive from a certain number
of uncertain constraints on. Algorithms based on gradient approximations of the second-
order cone are a newer development which can make use of the warm-start capabilities of
the simplex algorithm as the nodes in the branch-and-bound tree remain linear programs.
Furthermore, the quadratic description of the uncertain constraints is not fully incorpo-
rated into the problem from the start, but it is adaptively rened by gradient cuts. On the
other hand, this brings disadvantages of two kinds. In the context of robust optimization,
it requires many cutting planes to come to an adequate description of the ellipsoidal un-
certainty set, which makes nding robust feasible solutions harder. And in addition, the
LP relaxation obtained at the root of the branch-and-bound tree is not better than that
of the original nominal problem.

Therefore, we examine a dierent approach in this work, namely to linearize the ellip-
soidal uncertainty sets by applying the polyhedral approximation of the second-order cone
developed by Ben-Tal and Nemirovski (2001) and slightly improved by Glineur (2000).
Compared to a gradient approximation, it has the advantage of a compact representation
in terms of the number of variables and constraints. Their approach yields an outer ap-
proximation of the second-order cone. As a consequence, our polyhedral approximation of
an ellipsoidal uncertainty set will be an outer approximation, too. This ensures that every
scenario in the original ellipsoidal uncertainty set is considered in the arising approximate
robust counterpart. The advantage over other approaches using polyhedral uncertainty
sets  like the Γ-robustness proposed by Bertsimas and Sim (2004)  is that the original

© Springer Fachmedien Wiesbaden 2016


A. Bärmann, Solving Network Design Problems via Decomposition,
Aggregation and Approximation, DOI 10.1007/978-3-658-13913-1_11
160 10. Motivation

assumptions on the probability distribution of the uncertain coecients is retained, includ-


ing exact covariance information. Basically, our approach can be summarized as taking
the best of both worlds  the computational tractability of polyhedral uncertainty sets to-
gether with the more accurate description of the distributional information via ellipsoidal
uncertainty sets.

The typical way to obtain a compact representation of the robust counterpart to an un-
certain optimization problem is to introduce tight upper bounds on the security buer in
each row. These can be obtained by dualizing the corresponding subproblem which checks
for robust feasibility of a given solution. Here, we show how this can alternatively be done
by employing an inner approximation of the second-order cone. To this end, we derive a
new inner approximation based on the well-known outer approximation. The two behave
identically with respect to the size of the linear system and the accuracy guarantees. In
particular, our inner approximation is optimal with respect to the number of variables and
constraints.

We give computational results for our approach on two dierent types of problems. First,
we test it on portfolio optimization problems from the literature, which we use for a calabra-
tion of our method. Then we apply our approximation framework to optimal infrastructure
network design using nominal instances arising in railway network expansion. We will be
able to show that the proposed framework is eective in two respects. It can serve both
as a way to produce approximate solutions within a small margin of error in much shorter
computation time as the exact model and as a way to obtain high-quality solutions quickly
via a coarser approximation.

A more detailed analysis of our method and further computational experiments on adapted
MIPLIB and SNDlib instances as well as multidimensional-knapsack instances can be found
in Bärmann et al. (2015a). Together, our results show the eciency of the framework for
broad problem classes.

In the following, we give a classication of our approach in the context of the existing
literature. As the eld of robust optimization is broad, we only mention the works most
directly related to our study. For an introduction to robust optimization itself, we refer the
interested reader to Ben-Tal et al. (2009). Bertsimas et al. (2011) give a broad overview
of optimization under data uncertainty with a focus on robust optimization. In particular,
they review the tractability of the robust counterpart depending on the type of nominal
problem and the choice of the uncertainty set.

A model of data uncertainty that incorporates a budget of uncertainty for the coecients
of each constraint is introduced in Bertsimas and Sim (2004). It restricts the number
of coecients per row that can deviate from their nominal values, which is equivalent
to the choice of a specic polyhedral uncertainty set. This budget of uncertainty is less
conservative than protecting the model against all possible deviations of the coecients,
which would result in an axis-parallel box as the underlying uncertainty set. In Bertsimas
and Sim (2003), the authors focus on combinatorial optimization problems and especially
network ow problems. They derive a polynomial algorithm that solves the resulting robust
counterpart under budgeted uncertainty for polynomially solvable nominal combinatorial
problems. Here, a sequence of n+1 nominal problems is solved, where n is the number
10. Motivation 161

of binary variables. A similar result is obtained for problems for which a polynomial
approximation algorithm is known. The advantage of their approach, which is commonly
called Γ-robustness, is that the robust counterpart of a linear optimization problem remains
linear. However, we also lose exact covariance information.

Bertsimas and Sim (2006) develop a framework that uses approximations of the respective
uncertainty sets to maintain the problem class of the nominal optimization problem in
the robust counterpart. This is achieved by switching to uncertainty sets which are more
conservative but easier to handle than the those obtained from ellipsoidal distribution
assumptions. The disadvantage of this approach is that the idea is based on a weak relax-
ation, which means that the exact uncertainty set is contained in the relaxation without
an approximation guarantee.

An analysis over a set of NETLIB problems (see Gay (1985)) regarding the eect of data
uncertainty is performed in Ben-Tal and Nemirovski (2000). The observation is that such
uncertainty may lead to highly infeasible nominal solutions. Therefore, the level of con-
straint violation is determined, the nominal and the robust solutions are compared and
the so-called price of robustness is computed. As there is no given description of data un-
certainty for the NETLIB instances, the authors assume a certain perturbation set dened
as follows. A coecient is said to be uncertain if it cannot be represented as a rational
fraction
p
q, p ∈ Z, q ∈ N, with 1 ≤ q ≤ 100. It is shown that the choice of ellipsoidal
uncertainty sets lead to a high level of protection against data uncertainty. In Bärmann
et al. (2015a), we take similar assumptions for the evaluation of our framework on MIPLIB
instances.

Ben-Tal and Nemirovski (2001) introduce a compact outer approximation of the second-
order cone. Key to the approach is a decomposition of the n-dimensional second-order cone
Ln into copies of L2. These are in turn approximated via a homogenized approximation of
the unit disc B ⊂ R . The obtained approximation of B is an extended formulation of an
2 2 2

m-gon circumscribed to B . Its size in terms of the number of variables and constraints is
2

logarithmic in m, which is an advantage over the ordinary representation of the m-gon using
two variables and m constraints. It is also a potential advantage over an adaptive gradient
approximation as its cuts live in the same variable space. Furthermore, a lower bound of
O(n log 1 ) on the size of any polyhedral approximation of Ln with accuracy  is established.
In Glineur (2000), a slightly reduced representation of the extended formulation used for
the regular m-gon is given. Moreover, an optimal choice for m for each copy of L2 is
derived in order to obtain a prescribed accuracy for the approximation of Ln that achieves

the bound established in Ben-Tal and Nemirovski (2001).

Vielma et al. (2008) apply the extended formulation approach of Ben-Tal and Nemirovski
(2001) and Glineur (2000) to approximate conic quadratic constraints, which enables them
to tackle mixed-integer conic quadratic problems with linear solvers. This approach allows
for employing the warm-start capabilities of the simplex method in modern integer pro-
gramming solvers and thus has signicant advantages. In particular, it is shown that their
method outperforms commercial interior-point solvers as well as solvers based on gradient
approximations on portfolio optimization with integer variables.
162 10. Motivation

Our study can be seen as a thorough investigation of the eects observed in Vielma et al.
(2008). Motivated by those results, we build on the second-order cone approximation
of Ben-Tal and Nemirovski (2001) and Glineur (2000) to solve robust mixed-integer lin-
ear programs under ellipsoidal uncertainty. As indicated, staying within the realm of
mixed-integer linear programming enables us to leverage the full power of modern integer
programming solvers such as CPLEX and Gurobi.

To be able to derive upper bounds for the arising second-oder cone problems, we use the
second-order cone approximation for a conservative approximation of the ellipsoidal uncer-
tainty sets from outside. This allowed us to implement a generic automatic robustication
scheme which can be applied as an external library in combination with Gurobi. For a
given nominal problem and a description of the ellipsoidal uncertainty set, our program
formulates the SOC robust counterpart as well as its linear approximation. Additionally,
we develop a polyhedral inner approximation of the Lorentz cone from the polyhedral outer
approximation stated in Glineur (2000) and investigate the relation between the two. In
future implementations, the inner approximation will allow to avoid the tricky dualization
of the outer approximations of the uncertainty sets required to compute the security buer
in the robust counterpart.

After computational tests on the portfolio optimization instances from Vielma et al. (2008),
we present results on the original railway network instances from Part I. We emphasize
that our framework is not a new paradigm of robust optimization for network design
problems. Instead, it is way of solving the robust counterpart arising under the assumptions
of ellipsoidal uncertainty sets and static routing more quickly by approximation. For the
case of dynamic routing, we refer the reader to Poss and Raack (2013) and the references
therein. For alternative concepts of robust optimization in railway applications, such as
recoverable robustness and light robustness, we refer to the publications within project
ARRIVAL (2015), for example Liebchen et al. (2009) and Fischetti and Monaco (2009).

Our results conrm the superior performance of the extended-formulation based approxi-
mation approach observed in Vielma et al. (2008) for certain portfolio optimization prob-
lems. On other benchmark sets, it is vastly competitive and its performance is similar to
the gradient linearization approach used within Gurobi. Remarkable is that in all cases,
we obtain very good results with respect to the approximation quality. In particular, we
observe that our approach seems to be well suited to nd high-quality approximations
quickly in the case of the large-scale multi-commodity network design instances arising
from railway network expansion.
11. Polyhedral Approximation of
Second-Order Cone Robust
Counterparts

Our approach for an approximation of the second-order cone robust counterpart of an


uncertain mixed-integer program is based on a polyhedral approximation of the underlying
ellipsoidal uncertainty sets. These ellipsoidal uncertainty sets can be stated as the feasible
set of a second-order cone constraint intersected with some hyperplane. Therefore, we are
interested in a compact polyhedral approximation of the second-order cone. This chapter
starts by briey recalling the well-known outer approximation of the second-order cone by
Ben-Tal and Nemirovski (2001) and Glineur (2000). Then, building on their ideas, we show
how their construction can be used to derive an inner approximation of the second-order
cone which can be used to obtain upper bounds for second-order cone programs. Using
these bounds will allow us later to approximate the second-order cone robust counterpart
to an uncertain optimization problem. This new inner approximation behaves identically
to the well-known outer approximation of Ben-Tal and Nemirovski (2001) and Glineur
(2000) in terms of the size of the linear system and the accuracy guarantees. Especially, it
is possible to show that our inner approximation is optimal with respect to the number of
variables and constraints. The results on polyhedral approximations of the second-order
cone allow us to derive bounds for second-order cone optimization problems which we nally
use to obtain approximate robust counterparts with respect to ellipsoidal uncertainty.

11.1. Outer Approximation of the Second-Order Cone


We begin with the denition of the second-order cone, also called Lorentz cone:

Denition 11.1.1 (Second-order cone) . The second-order cone (SOC) Ln ⊂ Rn+1 is


dened as
Ln := {(r, x) ∈ R × Rn | kxk2 ≤ r}.
The rst step towards a polyhedral approximation of Ln is a decomposition into a linear
number of copies of L2 via an extended formulation. That means, we represent L as a
n

projection of a higher-dimensional representation using additional variables. Such a repre-


sentation is very benecial in many cases, as introducing few additional variables may allow
a drastic reduction in the number of constraints which are necessary to describe the desired
polyhedron. For an introduction to extended formulations, we refer the interested reader
to the excellent surveys by Conforti et al. (2010) and Kaibel (2011). The decomposition
of Ln into b n2 c copies of L2 is done as follows.

© Springer Fachmedien Wiesbaden 2016


A. Bärmann, Solving Network Design Problems via Decomposition,
Aggregation and Approximation, DOI 10.1007/978-3-658-13913-1_12
164 11. Polyhedral Approximation of Second-Order Cone Robust Counterparts

Theorem 11.1.2 (Ben-Tal and Nemirovski (2001)). The convex set


(
(r, x, y) ∈ R+ × Rn+b n
2
c
| x22i−1 + x22i ≤ yi2 , 1 ≤ i ≤ b n2 c,

L
( ))
dn2 (n even)
e
(r, y) ∈
(r, y, xn ) ∈ dn
L
2 (n odd)
e

is an extended formulation for Ln with b n2 c additional variables. The cone Ln is obtained


by projection onto the (r, x)-space.

Iteratively applying Theorem 11.1.2, we can decompose Ln into a total of n − 1 copies


of L2, using n − 2 additional variables. Each auxiliary variable appears in exactly two
dierent cones. Altogether, we obtain a recursion of logarithmic depth for the decomposi-
tion of Ln. This construction is sometimes referred to as the tower of variables. The term
guratively describes the structure of the auxiliary y -variables, which become fewer and
fewer in subsequent stages of the decomposition.

Now, in order to approximate Ln by a polyhedron, it suces to do so for each copy of L2 in


the decomposed representation of Ln. This can be done via a homogenized approximation
of the unit disc B . To quantify the accuracy of an outer approximation of B , we use the
2 2

notion of an outer -approximation. A polyhedron P is an outer -approximation of B if


2

B2 ⊆ P ⊆ {(1 + )x | x ∈ B2} holds. That means, an outer -approximation of B2 is a


polyhedron which contains B and which is itself contained in a slightly larger copy of B ,
2 2

namely B scaled up by (1 + ).


2

The natural choice for a polyhedral approximation B is the regular m-gon, and for the
2

sake of exposition, we normalize its rotation. Let Pm denote the (unique) regular m-gon
circumscribed to the unit disc B2 with one vertex at ( cos(1 π
m
) , 0). Given some  > 0, it can

be shown that any


 −1
1 π
m ≥ π arccos ≈√
+1 2
yields a regular m-gon Pm with B2 ⊂ Pm ⊂ (1 + )B2. The obvious way to describe Pm
would use two variables and m linear inequalities. Staying in the same variable space, it is
necessary to double the number of inequalities in order to increase the accuracy by a factor
of 4. For example, approximating B2 with an accuracy of 10−4 would already require 223
inequalities. Using an extended formulation, we can decrease the complexity drastically.

For the remainder of this thesis, let Z


γi := cos( 2πi ) and σi := sin( 2πi ) for i ∈ + . The following
theorem presents an extended formulation for Pm that only requires a logarithmic number
of dening inequalities relative to the level of accuracy.

Theorem 11.1.3 (Glineur (2000)). Let k ≥ 2. Then the polyhedron


 
 αi+1 = γ i α i + σ i βi (∀i = 0, . . . , k − 1) 
R
 
−βi+1 ≤ σ i α i − γ i βi (∀i = 0, . . . , k − 1)
 
:= 2k+2
Dk (α0 , . . . , αk , β0 , . . . , βk ) ∈

 −βi+1 ≤ −σi αi + γi βi (∀i = 0, . . . , k − 1) 

1 = γ k α k + σ k βk
 

is an extended formulation for P2k with projα0 ,β0 (Dk ) = P2k .


11.1 Outer Approximation of the Second-Order Cone 165

Remark 11.1.4. The extended formulation in Theorem 11.1.3 can be understood as a


deformed n-dimensional cube in Rn whose projection onto R2 yields a regular m-gon. We
refer the interested reader to Fiorini et al. (2012) and Kaibel and Pashkovich (2011) for
alternative constructions and interpretations.

The construction in Theorem 11.1.3 leads to a very ecient representation where the
π2
number of variables and inequalities is only logarithmic in m. Its accuracy is ≈ 22k+1
and
hence increasing k
1 (adding 2 variables, 1 equality and 2 inequalities) improves the
by
accuracy by a factor of 4. Continuing the example from above, an accuracy of 10−4 for the
approximation of B
2 can now be obtained for k = 8. The resulting extended formulation

has 18 variables, 9 equalities and 16 inequalities.

The notion of a polyhedral approximation of B2 naturally extends to Ln by homogeniza-


tion. A polyhedron L is an outer -approximation of Ln, if it satises Ln ⊆ L ⊆ {(r, x) ∈
R × Rn | kxk ≤ (1 + )r}. This leads to the following polyhedral approximation of L .
2

Theorem 11.1.5 (Glineur (2000)). The projection of the convex set


 
 αi+1 = γ i α i + σ i βi (∀i = 0, . . . , k − 1) 
R
 
−βi+1 ≤ −γi βi + σi αi (∀i = 0, . . . , k − 1)
 
L2 := (r, α0 , . . . , αk , β0 , . . . , βk ) ∈ 2k+3

 −βi+1 ≤ γ i βi − σ i α i (∀i = 0, . . . , k − 1) 

r = γ k α k + σ k βk
 

with  > 0 and k = dlog(π arccos( +1 ) )e onto the (r, α0 , β0 )-space is an outer -approxi-
1 −1

mation of L .
2

The representation of L2 can be reduced further by eliminating variables as done in Glineur
(2000). Combining the results from above, we obtain a polyhedral outer approximation of
the second-order cone Ln.
Theorem 11.1.6 (Glineur (2000)). Let Ln be the decomposition of Ln into n − 1 copies of
L obtained by the recursive application of Theorem 11.1.2. Approximating the L2-cones
2

according to Theorem 11.1.5 with accuracy l in stage l of the decomposition yields a set Ln
whose projection
Qt onto the (r, x)-coordinates is a polyhedral outer approximation of L with
n

accuracy  = k=1 (1 + l ) − 1, where t = dlog ne.

It is shown in Glineur (2000), Theorem 2.4, that the choice

   
π −1 l+1 16
l = cos( ) − 1, where ul = − log 4 ( ln(1 + )) ,
2u l 2 9π 2
in stage l of the decomposition leads to a polyhedral outer approximation Ln of Ln with
O(n log 1 ) variables and O(n log 1 ) inequalities if  < 12 . This implies the choice of a
u
homogenized 2 l -gon for the outer approximation of the L2-cones in stage l. It can be
shown that this is the minimal size of any polyhedral outer approximation of Ln with
prescribed accuracy  (see Ben-Tal and Nemirovski (2001)).
166 11. Polyhedral Approximation of Second-Order Cone Robust Counterparts

11.2. Inner Approximation of the Second-Order Cone


We now show how the construction of the outer approximation can be used to derive a
polyhedral inner approximation of Ln and characterize the relation between the two. The
practical benet of the inner approximation will be to provide a representation of the
dualized outer approximation within the linear approximation of the second-order cone
robust counterpart that is much easier to implement.

The inner approximation is also based on Theorem 11.1.2 for the recursive decomposition
of Ln into n − 1 copies of L2. However, L2 is now approximated via a homogenized inner
approximation of the unit disc B . An inner -approximation of B is a polyhedron P̄
2 2

satisfying
1
B 2 ⊆ P̄ ⊆ B2 . Similarly, a polyhedron L̄ is an inner -approximation of Ln

if it satises the inclusion {(r, x) ∈ R × R | kxk ≤


1+ r} ⊆ L̄ ⊆ L .
1+
n 1 n

The natural choice for a polyhedral inner -approximation of B is a regular m-gon inscribed
2

into B , which we denote by P̄m . Given some  > 0, it can easily be shown that any
2

m ≥ π arccos( 1+1 −1
) yields a regular m-gon P̄m with 1+ 1
B2 ⊆ P̄m ⊆ B2.
Remark 11.2.1. Observe the close relationship between an outer and an inner -approxi-
mation. Any given outer -approximation of B2 can be used to obtain an inner -approxi-
mation by inverting it at B2 . The same is true vice versa. This is illustrated in Figure 11.1
using the example of Pm and P̄m .

Figure 11.1.: The relationship between inner and outer m-gon approximations of B2

The approximation by regular m-gons is a special case of this relationship. Let m be big
enough such that Pm constitutes an outer -approximation of B2 . Then P̄m constitutes an
inner -approximation of B2 . Thus, it applies to both outer and inner approximation that
any
 −1
1
m ≥ π arccos
1+
suces to obtain an accuracy of .
In fact, for the regular m-gon, the outer and the inner approximation are polar to each
other. As before, it is not desirable to choose the obvious representation of P̄m with linearly
11.2 Inner Approximation of the Second-Order Cone 167

many constraints. A compact formulation of logarithmic size as developed in the following


π2
will again allow for an accurary of  ≈ 22k+1 using O(k) variables and constraints.

Remark 11.2.2 (Inner approximation from outer approximation via scaling) . Observe
that an extended formulation for an m-gon inscribed into B2 could also be obtained by
replacing
π π
1 = αk cos( k ) + βk sin( k ),
2 2
in the denition of the 2k -gon-approximation Dk for B2, with the equation
1 π π
= αk cos( k ) + βk sin( k ),
1 + k 2 2
with
1
k = − 1.
cos( 2πk )
The resulting polytope is an extended formulation of a 2k -gon P̄2k which satises
1
1 + k
B2 ⊂ P̄2 k ⊂ B2.
Consequently, replacing r = αk cos( 2πk ) + βk sin( 2πk ) analogously in the approximation L2 of
L2 yields an inner approximation of the latter. However, both polyhedral approximations
are prone to numerical instability, as they contain the coecient 1+ 1
k
≈ 1.

To avoid this instability (at least for moderate k ), we will now derive a direct inner ap-
proximation of B2 .
Theorem 11.2.3 (Bärmann et al. (2015a)). The polyhedron
 

 pi−1 = γ i p i + σ i di (∀i = 1, . . . , k − 1) 

−di−1 ≤ σi p i − γ i di (∀i = 1, . . . , k − 1)

 

 
R
 
di−1 ≤ σ i p i − γ i di (∀i = 1, . . . , k − 1)
 
2k
D̄k = (p0 , . . . , pk−1 , d0 , . . . , dk−1 ) ∈

 pk−1 = γk 

−dk−1 ≤ σk

 


 

dk−1 ≤ σk
 

for k ≥ 2 is an extended formulation for P̄2k with projp0 ,d0 (D̄k ) = P̄2k .

Proof. In the following, we describe the construction of the inner approximation as an


iterative procedure. We start by dening the polytope

Pk−1 := {(pk−1 , dk−1 ) | pk−1 = γk , −σk ≤ dk−1 ≤ σk }.

Now, we construct a sequence of polytopes Pk−1 , Pk−2 , . . . , P0 . Assume that polytope Pi


has already been constructed. In order to obtain polytope Pi−1 from polytope Pi , we
perform the following actions which we will translate into mathematical operations below:

π
1. Rotate Pi anticlockwise by an angle of θi = 2i
around the origin to obtain a poly-
tope Pi1 ,
168 11. Polyhedral Approximation of Second-Order Cone Robust Counterparts

Figure 11.2.: Construction of the inner approximation of the unit disc B2 for k = 3

2. Reect Pi1 at the x-axis to obtain a polytope Pi2 ,


3. Form the convex hull of Pi1 and Pi2 to obtain polytope Pi−1 .
The rst step is a simple rotation and can be represented by the linear map

R2 7→ R2,
    
x cos(θ) − sin(θ) x
Rθ : 7→ .
y sin(θ) cos(θ) y

The reection at the x-axis corresponds to the linear map

R2 7→ R2,
    
x 1 0 x
M: 7→ .
y 0 −1 y

Thus, the composition MRθi which rst applies Rθi and then M, is given by

R2 7→ R2,
    
x cos(θ) sin(θ) x
MRθi : 7→ .
y sin(θ) − cos(θ) y

With this, we obtain Pi1 = Rθi (Pi ) and Pi2 = (MRθi )(Pi ). Finally, adding the two
constraints
−di−1 ≤ σi pi − γi di
and
di−1 ≤ σi pi − γi di
yields a polyhedron whose projection onto the variables (di−1 , pi−1 ) is Pi−1 = conv(Pi1 , Pi2 ).
Keeping this correspondence in mind, we show that P0 = P̄2k .
In each iteration, Pi is rotated anticlockwise by an angle of θi around the origin, such
that the vertex of Pi with minimal vertical coordinate is rotated to (γk , σk ), therefore
Pi1 = R(Pi ). It is |V(Pi1 )| = |V(Pi )| and Pi1 lies strictly above the horizontal axis. Applying
11.2 Inner Approximation of the Second-Order Cone 169

the reection M, we obtain Pi2 = M(Pi1 ) which satises |V(Pi2 )| = |V(Pi1 )| and lies strictly
1 2
below the horizontal axis. Then Pi−1 = conv(Pi , Pi ) satises |V(Pi−1 )| = 2|V(Pi )| because
1 2
all vertices v ∈ V(Pi ) ∪ V(Pi ) remain extreme points of Pi . We obtain polytope P0 after
k − 1 iterations of the above procedure, which has |V(P0 )| = 2k vertices. As the interior
angles at each vertex of P0 are of equal size, it follows P0 = P̄2k . This proves the correctness
of our construction.

The intermediate steps of the construction are depicted in Figure 11.2 for the case k = 3,
which leads to an octagon approximation. The upper left picture shows the initial polytope
P2 , which is an interval on the line x = γk . The upper middle and upper right picture
show its rotation by 45◦ counterclockwise and clockwise, thus representing P21 and P22 ,
respectively. The lower left picture shows P1 as the convex hull of P21 and P22 . The lower
middle picture contains both P11 and P12 as a rotation of P1 by 90◦ in both directions.
Finally, the lower right picture shows P0 = P̄23 as the convex hull
1
of P1 and P1 .
2

Observe that the above proof can also be understood in terms of the framework of Kaibel
and Pashkovich (2011) using the notion of reection relations. From Theorem 11.2.3 follows
that the inner -approximation D̄k of B2 can be used to yield an inner -approximation of
L 2 by homogenizing it.

Corollary 11.2.4 (Bärmann et al. (2015a)) . The projection of the set


 

 pi−1 = γ i p i + σi di (∀i = 1, . . . , k − 1) 

−di−1 ≤ σ i p i − γ i di (∀i = 1, . . . , k − 1)

 

 
R
 
di−1 ≤ σ i p i − γ i di (∀i = 1, . . . , k − 1)
 
L̄2 = (s, p0 , . . . , pk−1 , d0 , . . . , dk−1 ) ∈ 2k+1

 pk−1 = γk s 

−dk−1 ≤ σk s

 


 

dk−1 ≤ σk s
 

with  > 0 and k = dlog(π arccos( +1 ) )e onto the variables (s, p0 , d0 ) is an inner -
1 −1

approximation of L .
2

This corollary allows for the derivation of an inner -approximation of Ln using a similar
construction as for the outer approximation presented in Section 11.1. Obviously, employ-
ing the decomposition technique from Theorem 11.1.2 and replacing the L2-cones by the
inner -approximation of the above corollary yields an inner approximation of Ln. Similar
to Theorem 11.1.6, we have

Theorem 11.2.5 (Bärmann et al. (2015a)). Let Ln be the decomposition of Ln into n − 1


copies of L as obtained by the recursive application of Theorem 11.1.2. Approximating
2

these L2 -cones according to Theorem 11.2.3 with accuracy l in the l-th stage of the decom-
position yields a set L̄n , whose projection
Q onto its (r, x)-coordinates is a polyhedral inner
approximation of Ln with accuracy  = tl=1 (1 + l ) − 1, where t = dlog ne.

Proof. To prove the accuracy claimed in the statement, it suces to show {(r, x) ∈ R × Rn |
kxk ≤ 1
1+ r} ⊆ projr,x (L̄n ), as the inclusion projr,x (L̄n ) ⊆ Ln is immediate.
170 11. Polyhedral Approximation of Second-Order Cone Robust Counterparts

Let (r, x) ∈ {(r, x) ∈ R × Rn | kxk ≤ 1+


1
r}, which is equivalent to k(1 + )xk ≤ r. Thus,
we have
n
X
r2 ≥ k(1 + )xk2 = (1 + )2 x2i .
i=1
With this in mind, we rst consider the case of even n. In the rst stage of the decom-
position, the condition (r, x) ∈ projr,x (L̄n ) requires us to nd suitable values of yi with
i = 1, . . . , n2 such that both

yi2 n
(x2i + x2i+1 ) ≤ for i = 1, . . . , (11.1)
(1 + 1 )2 2
and
n
2
X r2
yi2 ≤ (11.2)
(1 + 0 )2
i=1

hold. In Condition (11.2), 0 denotes the accuracy chosen for the approximation of the
remaining cone L n
2. Condition (11.1) can easily be satised by choosing yi ≥ 0 such that

n
yi2 = (1 + 1 )2 (x22i−1 + x22i ) for i = 1, . . . , .
2
We need to ensure that this choice also satises (11.2). For this observe

n n
2
X 2
X r2
yi2 = (1 + 1 )2 (x22i−1 + x22i ) = (1 + 1 )2 kxk2 ≤ (1 + 1 )2 .
(1 + )2
i=1 i=1

Thus, Condition (11.2) holds, if

(1 + 1 )2 1

(1 + )2 (1 + 0 )2
is satised, which in turn holds for

0
(1 + )2 ≥ (1 + 1 )2 (1 +  )2 .
0
This results in a guarantee of  = (1 + 1 )(1 +  ) − 1 for the accuracy of our inner
approximation of L n.

To attain the same guarantee for odd n, we need to show that our above choice for the yi 's
satises
n
b2c
X r2
x2n + yi2 ≤
(1 + 0 )2
i=1

instead of (11.2). This is possible by using

n n
b2c b2c n
X X X
x2n + yi2 = x2n + (1 + 1 )2 (x22i−1 + x22i ) ≤ (1 + 1 )2 x2i = (1 + 1 )2 kxk2 .
i=1 i=1 i=1
0
From here, we can proceed as in the even case to obtain a guarantee of  = (1+1 )(1+ )−1.
11.3 Upper Bounds for Second-Order Cone Programs via Linear Approximation 171

The proof for the rst stage of the decomposition can be applied recursively to the remain-
ing cone in each subsequent stage. Altogether, we have shown that the accuracy of the
Qt
inner approximation is = l=1 (1 + i ) − 1, where t = dlog ne is the number of stages of
the decomposition.

Remark 11.2.6. The above proof indicates that there is no obvious gain in accuracy by
choosing dierent accuracies for the L2 -cones across a given stage of the decomposition.
For accuracies i1 , i = 1, . . . , b n2 c, in stage 1, we estimate
n n
b2c b2c
X X
yi2 = (1 + i1 )2 (x22i−1 + x22i ) ≤ (1 + max i1 )2 kxk2 .
i=1,...,b n
2
c
i=1 i=1

This can be made reasonably tight, so that the overall accuracy in a given stage mainly
depends on the worst accuracy of any L2 -approximation therein.

Having established a guarantee for the quality of the inner approximation of Ln, the
remaining question is how to choose the values of l in stage l of the decomposition in
order to obtain a representation of small size. This can be answered using Remark 11.2.1
if we keep in mind from the proof of Theorem 11.2.5 that outer and inner approximation
of Ln obey the same formula for the dependence of the overall accuracy. It follows that
the size of our inner -approximation of L is both O(n log ) in the number of variables
n 1

1
and constraints if < 2.

Corollary 11.2.7 .
Choosing intermediate accuracies
(Bärmann et al. (2015a))
   
π l+1 16
l = cos( u )−1 − 1, where ul = − log4 ( 2 ln(1 + )) ,
2 l 2 9π
for the inner L2 -approximations in stage l of the decomposition allows for an inner ap-
proximation of Ln using O(n log 1 ) variables and O(n log 1 ) constraints if  < 12 .

Regardless of potential numerical issues, scaling our inner -approximation for B2 by


(1 + ) yields an outer approximation of 2 . Hence, B the construction yields an outer
L
-approximation of n and is therefore subject to the lower bound established in Ben-
Tal and Nemirovski (2001), showing that it is optimal in terms of the size of the linear
system.

Corollary 11.2.8 (Bärmann et al. (2015a)). The size of the inner -approximation of Ln
stated in Theorem 11.2.5 using the intermediate accuracies of Corollary 11.2.7 is optimal
in the sense that there is no other inner -approximation of Ln using fewer variables and
constraints asymptotically.

11.3. Upper Bounds for Second-Order Cone Programs via


Linear Approximation
In order to derive a linear approximation for the SOC robust counterpart of an LP or MIP,
we need to be able to nd tight upper bounds for SOCPs. Therefore, we consider the
172 11. Polyhedral Approximation of Second-Order Cone Robust Counterparts

following second-order cone program (P) with a single second-order cone constraint:
 
r
max cT
x
 
r
s.t. A +b = 0
x
(r, x) ∈ Ln
for a matrix A∈ Rm×(n+1) and vectors b ∈ Rm and c ∈ Rn+1. A straightforward upper
bound can be obtained by solving the associated dual problem (D):

min bT y
 
ρ
s.t. AT y − +c = 0
ξ
(ρ, ξ) ∈ Ln .
This primal-dual pair has several convenient properties, of which the following strong
duality result is the most important one for us:

Theorem 11.3.1 (Glineur (2001)). If both (P) and (D) possess a strictly feasible solution,
both problems attain their respective optimal values p̄ and d¯ and the latter two coincide,
i.e. there are optimal solutions (r̄, x̄) to (P) and (ȳ, ρ̄, ξ)
¯ to (D) with
 

cT = p̄ = d¯ = bT ȳ.

The above theorem allows us to substitute the SOC maximization subproblem of the
form (P) as it arises within the SOC robust counterpart by the corresponding dual (D),
retaining the same optimal value. This enables us to take two dierent ways to obtain
tight upper bounds for the optimal value of (P) by means of polyhedral approximation.
The rst is to replace the constraint (r, x) ∈ Ln in (P) by an outer approximation, the
second is to replace the constraint (ρ, ξ) ∈ n L in (D) by an inner approximation. We will
make use of this property when approximating the auxiliary optimization problem which
nds a worst-case scenario from an ellipsoidal uncertainty set to a given solution to the
nominal problem. It will allows us to derive compact linearizations of the SOC robust
counterpart.

11.4. Approximating SOC Robust Counterparts


The results on polyhedral approximations of the second-order cone from the previous chap-
ter will now be applied to derive approximate robust counterparts under ellipsoidal uncer-
tainty. We begin with a short revision of the basic concept of robust optimization. After
stating the robust counterpart under polyhedral and ellipsoidal uncertainty respectively,
we derive compact linear approximations of the SOC robust counterpart.
11.4 Approximating SOC Robust Counterparts 173

11.4.1. Basics in Robust Optimization


Consider the following mixed-integer linear optimization problem:

min cT x
s.t. Âx ≤ b (11.3)

x ∈ Z
p
+ × R n−p
+

for a matrix R R R
 ∈ m×n , vectors c ∈ n and b ∈ m and a number of integer variables
1 ≤ p ≤ n. We call (11.3) the nominal optimization problem for the nominal realization
 of the uncertain constraint matrix A. If this matrix A is subject to data uncertainty
according to a closed uncertainty set U , its robust counterpart is dened as

min cT x
s.t. Ax ≤ b (∀A ∈ U ) (11.4)

x ∈ Z p
+ × R n−p
+ .

This statement of the robust counterpart possesses one set of constraints Ax ≤ b for each
possible realization of A. A vector x ∈ Rn which is feasible for (11.4) is called robust
feasible. Especially, such a robust feasible solution is feasible for all possible realizations of
the uncertain matrix A.
Given the nominal matrix Â, the uncertainty set U can be stated as

U = {A ∈ Rm×n | A = Â + Ã, Ã ∈ S}, (11.5)

where S is a closed set, the so-called perturbation set. Assuming this form of the uncertainty
set U , we can give a more intuitive formulation of the robust counterpart. In the following,
let Ai denote the i-th row of a matrix A.
Theorem 11.4.1 (Bertsimas et al. (2004)). The robust counterpart for a closed uncertainty
set U given in the form (11.5) can be stated as

min cT x
s.t. Âi x + max Ãi x ≤ bi (∀i = 1, . . . , m) (11.6)
Ã∈S
x ∈ Zp+ × Rn−p
+ .

This characterization suggests the following interpretation of the robust counterpart: Each
constraint of the nominal problem is immunized against uncertainty in the data by adding
a security buer. The latter is chosen according to the worst possible outcome of the data
realization for that specic row. Computing the respective security buer of each row for
a given solution candidate x̂ ∈ Rn amounts to solving the following family of subproblems
for i = 1, . . . , m:
max Ãi x̂
(11.7)
s.t. Ã ∈ S.
174 11. Polyhedral Approximation of Second-Order Cone Robust Counterparts

Thus, it suces to consider constraint-wise uncertainty, i.e. the case of a separate uncer-
tainty set Ui for each row Ai of the constraint matrix A. For the sake of exposition and
without loss of generality, we therefore conne ourselves to nominal problems with one
uncertain constraint a T x ≤ b, a ∈ Rn and b ∈ R, for the rest of this chapter.

11.4.2. The Robust Counterpart under Linear and Ellipsoidal Uncertainty

For a concrete uncertainty set U , it is possible to state the corresponding robust counterpart
in a closed form if there exists a dual problem to (11.7) such that strong duality holds.
This is the case for polyhedral and ellipsoidal uncertainty sets.

We rst consider the case where the perturbation set S is given as a polyhedron, i.e.

U = {a ∈ Rn | a = â + ã, Dã ≤ d}, (11.8)

where D is a matrix in Rr×n and d is a vector in Rr . In this case, the robust counterpart
can be formulated as a linear program again.

Theorem 11.4.2 (Bertsimas et al. (2011)). The robust counterpart (11.6) to the nominal
problem (11.3) with â ∈ U and U as in (11.8) can be written as:

min cT x
s.t. âT x + dT π ≤ b
DT π = x (11.9)

x ∈ Z p
+ × R n−p
+

π ≥ 0.

The above also holds for mixed-integer problems as long as we consider continuously valued
data uncertainty as the derivation only requires strong duality for Subproblem (11.7). The
x-variables are xed here, and therefore (11.7) only consists of continuous variables.

The aim of our computational study is to analyse the performance of solving robust mixed-
integer optimization problems with ellipsoidal uncertainty sets where the latter are approx-
imated via polyhedral approximation of the second-order cone. The choice of ellipsoidal
uncertainty sets is motivated by the approximation of so-called probabilistic or chance con-
straints as well as mean-variance optimization problems (see Ben-Tal et al. (2009) for a
detailed treatment). We consider the robust counterpart for an ellipsoidal uncertainty set
U of the form:

U = {a ∈ Rn | a = â + Bz, kzk ≤ 1}. (11.10)

The corresponding robust counterpart is a conic quadratic optimization problem.


11.4 Approximating SOC Robust Counterparts 175

Theorem 11.4.3 (Ben-Tal and Nemirovski (2000)). The robust counterpart (11.6) to the
nominal problem (11.3) with â ∈ U and U as in (11.10) can be written as
min cT x
s.t. âT x + kB T xk ≤ b (11.11)

x ∈ Z p
+ × R
n−p
+ .

We see that optimizing the robust counterpart (11.6) under ellipsoidal uncertainty involves
the solution of a second-order cone problem. Especially in the presence of integral variables,
genuinely quadratic methods for the solution of such problems, like interior point methods
(see for example Bonnans et al. (2009)), suer from the lack of warm-start capabilities.
Thus, in a branch-and-bound framework, they cannot reuse the optimal solution to a node
in the branch-and-bound tree in the solution of the child nodes, which is a signicant
disadvantage. In contrast to that, the robust counterpart under polyhedral uncertainty
is a mixed-integer linear program. As such, it can make ecient use of warm-start as
well as other advanced MIP techniques like cutting planes and preprocessing, which are
implemented in standard MIP solvers.

11.4.3. Linear Approximation of the SOC Robust Counterpart


We would like to benet from the computational advantages of polyhedral uncertainty
sets while retaining the stochastic information of the original ellipsoidal uncertainty set
at the same time. To this end, we derive a linear approximation of the second-order
cone robust counterpart (11.11) which exploits the structure of the ellipsoidal uncertainty
set. Replacing the unit-ball Bn in the denition of U by an outer -approximation B n
obviously yields an outer -approximation of U of U. This is advantageous as U contains
all the scenarios in U, and therefore the approximate robust counterpart according to
U is guaranteed to produce feasible robust solutions to the original robust counterpart
according to U. On the other hand, an outer -approximation ensures that U is not overly
conservative, as its scenarios are either contained in U or at least very close to other
scenarios in U.
Let U be given as in (11.10). Substituting z = B −1 (a − â) yields

Rn | kB −1(a − â)k ≤ 1} = {a ∈ Rn | B −1(a − â) ∈ Bn}.


U = {a ∈
Now let B be a polyhedral outer -approximation of B given by B = {z | Kz ≤ k}.
n n n

Note that this notation assumes the outer approximation to live in the same space of
variables as the unit ball itself, which would be the situation after applying projection to the
outer approximation from Section 11.1. However, transferring the following considerations
to the case of an approximation given by an extended formulation is straightforward.
Replacement of Bn by Bn in the last equation yields the approximate uncertainty set U ,
which can be stated as follows:

U = {a ∈ Rn | B −1(a − â) ∈ Bn}


= {a ∈ Rn | K(B −1(a − â)) ≤ k}
= {a ∈ Rn | KB −1a ≤ k + KB −1â}.
176 11. Polyhedral Approximation of Second-Order Cone Robust Counterparts

Thus, for the choice D := KB −1 and d := k + KB −1 â, the system Da ≤ d denes


the polyhedron U approximating U . Using Theorem 11.4.2, we can state the following
formulation of the corresponding approximate robust counterpart.

Corollary 11.4.4 (Bärmann et al. (2015a)). Let the uncertainty set U be given by U =
R B
{a ∈ n | B −1 (a − â) ∈ B n }, where B n is an outer -approximation of n with B n = {z |
Kz ≤ k}. Then the robust counterpart (11.9) takes the form
min cT x
s.t. âT x + k T π ≤ b
BT x − K T π = 0
x ∈ Zp+ × Rn−p
+

π ≥ 0.

Building on the discussion in Section 11.3, it as well possible to state the approximate
robust counterpart in terms of an inner approximation of Ln . To see this, consider Sub-
problem (11.7), which returns a worst-case scenario from the set U for a given solution x̂ to
the original problem (11.3). For ellipsoidal U as in (11.10), it is an SOCP of the following
form:
max x̂T a
(11.12)
s.t. B −1 a ∈ Bn .
Its dual can be written as:

min ρ
(11.13)
s.t. (ρ, B T x̂) ∈ Ln.
Now, using Theorem 11.3.1, we obtain an upper bound for Subproblem (11.12) by replacing
Ln in the dual problem (11.13) by an inner approximation. Let this inner approximation
be given by the polyhedral set L̄n = {(ρ, ξ) | W ξ ≤ wρ}, again assuming a representation
in the original variable space. The approximate subproblem then reads:

min ρ
(11.14)
s.t. W B T x̂ ≤ wρ.

This allows us again to state the approximate linear robust counterpart in a closed form:

Theorem 11.4.5 (Bärmann et al. (2015a)). Let the uncertainty set U be given by U =
R B
{a ∈ n | B −1 (a − â) ∈ n }, and let L̄n = {(ρ, ξ) | W ξ ≤ wρ} be an inner approximation
L
of n . Then the projection of the feasible set of the optimization problem
min cT x
s.t. âT x + ρ ≤ b
W B T x ≤ wρ
x ∈ Zp+ × Rn−p
+
11.4 Approximating SOC Robust Counterparts 177

onto variable x is a subset of the feasible set of the SOC robust counterpart (11.11).

The immediate consequence of the above reasoning is that we can establish a guarantee
of factor (1 + ) on the approximation of the security buer of an uncertain constraint,
as this is the optimal value of the approximate subproblem (11.14) if Ln is approximated
with accuracy  from inside. For an outer -approximation, the same bound follows from
inspection of the corresponding subproblem 11.7.

Altogether, we see that a tight outer approximation of the uncertainty sets yields a tight
upper bound on the security buer in the approximately robustied constraint. As a
consequence, we may hope to obtain a tight approximation of the optimal robust objective
value, too, although this is not guaranteed, especially in the presence of integral variables.
The following computational results include an empirical study on the inuence of the
chosen accuracy  on the quality of the approximate robust objective value which justies
this hope.
12. Computational Assessment of
Approximate Robust Optimization

In this chapter, we present the results of our approximate robust optimization framework
for dierent benchmarks sets. We compare them to the results obtained for the exact
quadratic formulation of the robust problems with respect to dierent performance mea-
sures. Special focus is put on the quality of the approximate solution in relation to the time
needed to compute it. We will see that solutions of very high quality are generally available
much faster by our approximation scheme than by solution of the quadratic model.

We begin with a description of implementation details and the computational setup before
we describe the instances of the benchmark sets. These are portfolio optimization instances
and instances from railway network expansion. Then we give a detailed presentation and
analysis of the computational results.

12.1. Implementation and Test Instances


We have implemented an exact quadratic robust optimization framework as well as our
approximate framework as described previously. The resulting command line tool takes
as input the nominal problem in any standard format for linear programs together with a
le stating the uncertain parameters and the matrices dening their ellipsoidal uncertainty
sets. It then solves either the quadratic or the approximating linear robust counterpart,
as specied. For the implementation, we used the C++-Interface of Gurobi 5.0. As an
alternative to this command line tool, it is also possible to link our modeling framework to
a given source code in C++ as a library in order to implement robust optimization models
via Gurobi.

Our implementation uses the formulation given in Corollary 11.4.4 building on dualized
outer approximations of the second-order cones as this was the most obvious way to go
when starting our work. In the process we noted that an implementation in terms of
Theorem 11.4.5 using inner approximations would be much easier to accomplish as it
would avoid the complicated dualization of the nested decomposition according to Theo-
rem 11.1.2. Rather than rst applying Theorem 11.1.2 to the outer approximations of the
L2-cones and then dualizing this tower-of-variables construction, we can directly use the
primal representation of the decomposition together with inner approximations of the L -
2

cones. This will come in handy in future implementations. Our discussion in the previous
sections already showed that the approximation guarantee is the same in both cases, which
also applies to the size of the arising linear system.

© Springer Fachmedien Wiesbaden 2016


A. Bärmann, Solving Network Design Problems via Decomposition,
Aggregation and Approximation, DOI 10.1007/978-3-658-13913-1_13
180 12. Computational Assessment of Approximate Robust Optimization

In the following, we compare the exact quadratic model and our approximate polyhedral
approach on two dierent types of test instances. These are the portfolio optimization
instances used in Vielma et al. (2008), which we use for calibration purposes, and network
design instances from railway network expansion. Our results on further benchmark sets
can be found in Bärmann et al. (2015a).

In the tables and diagrams presented here, QR generally refers to results for the solution of
the quadratic robust counterpart while LR- stands for the linearized robust counterpart
with an approximation accuracy of . For the latter, we will particularly focus on the
quality of the approximation as well as its computational complexity. The computations
were performed on a compute server with Six-Core AMD Opteron— 2435 processors and
64 GB RAM using all six cores and a time limit of one hour per instance. Unless stated
otherwise, we used Gurobi's default parameter settings.

12.2. Results on Portfolio Instances from the Literature


We rst consider a set of portfolio instances from Vielma et al. (2008), which were used to
test their algorithm for mixed integer conic quadratic programming problems. We apply
our approach to these instances originating from real-world stock market data and use them
to calibrate the parameters of our method. Our performance results are vastly consistent
with those presented in Vielma et al. (2008).

12.2.1. Description of the Portfolio Instances


Vielma et al. (2008) test their algorithm on three dierent versions of the portfolio opti-
mization problem, which they denote by classical, shortfall and robust instances. The rst
version  classical  maximizes the expected return of the portfolio while keeping its risk
below a certain predened threshold. The shortfall instances dier in the risk measure
used, while the robust instances consider the returns as uncertain. For our benchmark,
we use the classical problem version, which is formulated as the following mathematical
program:
max aT y
s.t. kQ1/2 yk2 ≤ σ
yi ≤ x i (∀i ∈ {1, . . . , n})
n
P
xi ≤ K (12.1)
i=1
n
P
yi = 1
i=1
x ∈ {0, 1}n
y ∈ [0, 1]n ,

where R
a ∈ n is the vector of expected returns for a given number n of available assets.
Matrix Q1/2 ∈ Rn×n is the positive semidenite matrix square root of the covariance
12.2 Results on Portfolio Instances from the Literature 181

matrix of these returns. The aim is then to invest into at most K<n dierent assets to
maximize the return of the entire portfolio while keeping its risk below a level of σ.
Variable xi indicates whether asset i = 1, . . . , n is chosen to invest into, while variable yi
indicates the fraction of the total investment allocated to that asset. The rst constraint
restricts the risk of the investment to ensure that it is not higher than σ. It is formulated
as a second-order cone constraint. The other constraints limit the total number of dierent
assets to be purchased to at most K and enforce the investment of the entire budget into
the portfolio.

The left-hand side of the second-order cone risk constraint possesses an interpretation in
the context of robust optimization. It can be seen as the security term arising as the
optimal value of Subproblem (11.12), which ensures that the solution remains feasible
under all possible scenarios within a given ellipsoidal uncertainty set. The latter is dened
by covariance matrix Q. In this interpretation, Problem (12.1) is the corresponding SOC
robust counterpart (11.11). To t the problem into our computational framework, we use
the Cholesky root of the covariance matrix instead of the matrix square root, which in
theory makes no dierence for the evaluation of the norm term.

In Vielma et al. (2008), results for instances with n ∈ {20, 30, 50, 100, 200} were reported.
As the case n = 30 is vastly similar to the case n = 20, and because the instances for
n = 200 were still too dicult to solve within the time limit by any of the methods, we
present results for their complete instance set for n ∈ {20, 50, 100}. These are 100 instances
each for the rst two cases and 10 instances for n = 100.

12.2.2. Results on Small Instances: Calibration of Parameters


In the following, we summarize our results on instances with n = 20 assets, which we use to
calibrate our framework, and describe the parameter settings arising from the observations
on these instances.

Solution Times To get a proper idea of the complexity of these easier instances, we begin
by stating the solution times for all algorithms under consideration in Table 12.1.

QR LR

time [s] IP GL =1  = 0.1  = 0.05  = 0.01


min 0.10 0.03 0.01 0.01 0.04 0.02
max 2.38 28.38 0.49 0.36 0.41 0.53
mean 0.43 2.05 0.14 0.15 0.19 0.23

Table 12.1.: Solution times [s] for instances with 20 assets

Gurobi has two built-in algorithms which can be applied to solve the SOC robust coun-
terpart, whose solution times are found under QR. The rst is its interior-point solver
182 12. Computational Assessment of Approximate Robust Optimization

and the second its gradient linearization algorithm, denoted by IP and GL respectively.
The abbreviation LR refers to our linear relaxation of the uncertainty set to approximate
the SOC robust counterpart. We used four dierent approximation accuracies, namely
 ∈ {0.01, 0.05, 0.1, 1}.

It can be seen that this subset of instances does not present a big problem for any of the
algorithms, although certain dierences are obvious. The interior-point algorithm performs
much better than the gradient linearization. The average solution time of the former is 5
times faster, while in the worst case the gradient linearization takes 14 times longer in the
worst case.

This is supported by the performance prole in Figure 12.1, which gives a visualization of
the solution times of the two Gurobi algorithms.

100

90

80

70
% of instances solved

60

50

40

30

20

10
QR-IP
QR-GL
0
1 10 100
Multiple of fastest solution time (log-scale)

Figure 12.1.: Performance prole for the instances with 20 assets and the two quadratic
algorithms

It becomes visible that the IP solver was faster on more than 70 % of all instances, and
that the gradient linearization does not catch up even when allowing higher multiples of
the fastest solution time.

However, from Table 12.1 it also becomes obvious that both Gurobi algorithms are dom-
inated by any of our linear approximations of the robust counterpart. The mean solution
time of even the nest approximation is only half of that of the IP solver while its maximal
solution time is at most one fourth as long. We note in advance that the trend visible in
the increasing mean solution time for a ner approximation becomes much more distinctive
for many of the harder benchmark instances presented here.
12.2 Results on Portfolio Instances from the Literature 183

Parameter Settings As shown in Theorem 11.2.5, we can obtain a polyhedral approxima-


tion of the second-order cone with an a priori approximation accuracy of , which consists
of O(n log 1 ) variables and inequalities for dimension n. Obviously, there is a trade-o
between a tighter polyhedral relaxation for smaller values of  and a smaller linear system
describing the approximation for bigger values of . Therefore, we want to determine an
 which ensures good results in terms of small approximation errors while maintaining
reasonable problem sizes, which has a major impact on computation times.

In Table 12.2, we show the minimum, maximum and average error as well as the standard
deviation for optimal solution of the polyhedral approximation in comparison to the exact
optimal value for the instances with n = 20. The error measured in per cent is dened
as
|zQR − zLR- |
· 100,
|zQR |

where zQR and zLR- are the optimal value of the quadratic model and the optimal value
of our linear approximation with acccuracy  respectively.

error [%]  = 0.01  = 0.05  = 0.1 =1


min 0.11 0.22 0.81 2.46
max 0.27 0.74 2.81 10.67
mean 0.19 0.48 1.70 6.80
std 0.03 0.10 0.44 1.66

Table 12.2.: Approximation error [%] for dierent a priori accuracies  for 20 assets

As expected, we see that the approximation error decreases monotonously in . These


computations suggest to choose  = 0.01 or  = 0.05 as both yield approximation errors
below 0.5 %. A performance plot of the solution times of the approximations is given in
Figure 12.2. It is consistent with the gures in the Table 12.1 and shows a clear solution
time advantage for  = 0.05 to  = 0.01. This dominance becomes more evident for larger
instances  up to the point of non-competitiveness of the latter. Therefore, we choose
 = 0.05 for our experiments on the harder portfolio instances in order to show that near-
optimal solutions can be found in much less computation time. For our results on large-scale
railway network expansion instances, we will use the coarsest possible approximation to
indicate the eciency of our method as a quick heuristic which already yields high-quality
results.

To cope with the size of the linear system arising for  = 0.05 in larger instances, the root
node of the branch-and-bound tree is solved using the barrier algorithm in the case of the
linear approximation. All later nodes are solved via the dual simplex algorithm in order
to prot from its warm-start capabilities. The two Gurobi algorithms are always run with
standard parameter settings.
184 12. Computational Assessment of Approximate Robust Optimization

100

90

80

70
% of instances solved

60

50

40

30

20

LR-1
10 LR-0.1
LR-0.05
LR-0.01
0
1 10 100
Multiple of fastest solution time (log-scale)

Figure 12.2.: Performance prole for 20 assets and varying values of 

12.2.3. Results for Larger Instances

The results presented in the following are for larger portfolio instances with n = 50 and
n = 100 assets. We compare the performance of an approximation with  = 0.05 to that
of exact solution via Gurobi's IP solver in Tables 12.3 and 12.4.

QR LR QR LR

time [s] IP GL  = 0.05 time [s] IP GL  = 0.05


min 0.35 ∞ 0.31 min 141.11 ∞ 88.89
max 2048.11 ∞ 581.50 max ∞ ∞ ∞
mean 91.59 ∞ 16.42 mean 2375.53 ∞ 1056.66

Table 12.3.: Solution times for 50 assets Table 12.4.: Solution times for 100 assets

It becomes visible at rst sight that Gurobi's gradient linearization is not competitive on
these instances as it could not solve any of them within the time limit. This is in sharp
contrast to our other results published in Bärmann et al. (2015a) on other benchmark sets,
where it was the other way round. A reason for this behaviour might be that the portfolio
problem involves the robustication of an otherwise empty row of the problem. This
might lead to a high number of gradient cuts which are necessary to achieve a tight linear
description of the quadratic constraint.
12.3 Results on Real-World Railway Instances 185

For n = 50 assets, both the IP solver and our linear approximation with accuracy  = 0.05
could solve all of the instances to optimality. The average solution time of our approxi-
mation is more than 5 times faster than the time required by the exact IP solver. For the
largest instances in this benchmark set, n = 100 assets, neither method could solve all the
instances. The linear approximation still prevails on these instances by a factor of two.

Our results on the portfolio instances are rounded o with two more performance plots in
Figures 12.3 and 12.4. They exclusively compare the IP solver to the linear approximation
in the cases of n = 50 and n = 100 assets respectively.

100 100

90 90

80 80

70 70
% of instances solved

% of instances solved
60 60

50 50

40 40

30 30

20 20

10 10
QR-IP QR-IP
LR-0.05 LR-0.05
0 0
1 10 100 1 10 100
Multiple of fastest solution time (log-scale) Multiple of fastest solution time (log-scale)

Figure 12.3.: Performance prole com- Figure 12.4.: Performance prole com-
paring Gurobi's interior-point solver and paring Gurobi's interior-point solver and
a linear approximation with  = 0.05 on a linear approximation with  = 0.05 on
portfolio instances with 50 assets portfolio instances with 100 assets

In the case n = 50, almost 95 % of all instances were solved fastest by our approximation
and it took large multiples of time for the IP algorithm to catch up. In the case n =
100, all solvable instances were solved fastest by our method. The conclusion on this set
of benchmark instances is that the linear approximation is well suited to obtain near-
optimal solutions within relatively short time in the context of portfolio-like optimization
problems.

12.3. Results on Real-World Railway Instances


We will now present the results of our approximation framework on two sets of railway
network expansion instances which are derived from the instances considered in Chapter 5
from Part I. More precisely, we consider the single-period expansion problem (NEP) for
the eight railway networks from Table 5.5 on page 93 and assume a data uncertainty in the
demand gures. For each demand relation, we assume that the actual number of trains
to be routed is within ±3 the nominal number along all paths which are at most 2 %
longer than the shortest path for this relation. For this choice of the uncertain parameters,
we incorporate a robust version of the capacity constraints of the problem. The shortest
paths in the network as well as short detours are most likely to become congested if the
demand gures are somewhat higher than forecasted. Therefore, our assumptions can be
186 12. Computational Assessment of Approximate Robust Optimization

interpreted as a local robustication of the network design by creating security buers. The
chosen magnitude of the demand uncertainty corresponds to a deviation of about 10 % for
a somewhat higher frequented relation.

The two benchmark sets considered here dier in the number of relations which are re-
garded as uncertain. In the rst experiment, we want to examine if our approach is
signicantly faster in computing near-optimal robust solutions for a moderate number of
uncertain coecients than the exact quadratic model. To this end, we consider the rst 10
coecients of each capacity constraint to be uncertain and solve the instances with a high
approximation quality, namely for  = 0.05. In the second case, we test the capability of
our method to produce high-quality solutions for a high number of uncertain coecients
within short computation time. This is done for the coarsest possible approximation,  = 1,
considering the rst 200 coecients in each capacity constraint to be uncertain. In both
cases, the root relaxation of the approximation is solved via the dual simplex method, as
this showed to be more ecient in preliminary results. Again, all computations were given
a time limit of 1 h.

12.3.1. Results under Moderate Uncertainty

We begin by assessing the suitability of our method as an actual approximation frame-


work, i.e. the possibility to obtain near-optimal solutions faster than computing the exact
optimum. Table 12.5 states the results for up to 10 uncertain coecients in the capacity
constraints comparing our compact polyhedral approximation with  = 0.05 to Gurobi's
gradient-based linearization. As already indicated, Gurobi's interior-point algorithm does
not seem to be the algorithm of choice in the presence of many uncertain constraints.
Accordingly, its performance was vastly uncompetitive in preliminary experiments on the
instances presented here, which is why we leave out the details.

LR-0.05 QR-GL
Solution time or Solution quality Solution time or
Instance Remaining gap plain / adjusted Remaining gap

SH-HA-12 2s < 0.01 % / 0.04 % 5s


B-BB-MV-5 7s 0.00 % / 0.00 % 39 s
NS-BR-180 90 s < 0.01 % / 0.01 % 127 s
HE-RP-SL-295 12 s < 0.01 % / < 0.01 % 66 s
TH-S-SA-60 421 s 0.01 % / 0.41 % 1701 s
BA-BW-165 2324 s < 0.01 % / 0.06 % 1560 s
NRW-90 51 s < 0.01 % / < 0.01 % 516 s
Deutschland-700 2590 s − −

Table 12.5.: Performance of a tight polyhedral approximation compared to Gurobi's gra-


dient linearization on railway instances with moderate uncertainty
12.3 Results on Real-World Railway Instances 187

The results in Table 12.5 show a very favourable behaviour of the polyhedral approxima-
tion. First of all, we see that the approximation can be solved signicantly faster than the
exact quadratic model for almost all the instances. As example, the approximate robust
counterpart for instance NRW-90 is solved ten times faster than the exact one. Moreover,
instance Deutschland-700 cannot be solved exactly at all as gradient linearization does not
even nd a feasible solution, while the approximation is available within the given time
limit.

Complementing the solution times, we also investigated the quality of the obtained approx-
imate solutions. Similar to our approach in Section 5.3.1 from Part I we state a plain
and an adjusted measure of quality. Recall that we introduced this distinction in order
to account for the prots of the railway undertakings that are possible even without up-
grading the network. The plain measure is a direct comparison of the objective values,
while the adjusted measure evaluates the actual improvement of the routing by implement-
ing the chosen upgrades. Here we see that the plain objective value of the approximate
solution is always within 0.01 % of the exact optimal solution value (except for the case
of Deutschland-700 where no such comparative value is available). Passing to the ajusted
measure, we see that the approximate solution is at most half a per cent worse than the
exact solution. This shows that a tight approximation can be obtained much faster than
the exact solution. Finally, considering instance BA-BW-165, whose approximation took
somewhat more time than the exact solution, we point out that the approximate robust
counterpart does not necessarily have to be solved to optimality in order to obtain a so-
lution of high quality. The approximation scheme only required 235 s to nd a solution
within 1 % (adjusted), which is considerably shorter than the 1701 s required for an optimal
solution.

12.3.2. Results under High Uncertainty

In our second benchmark set, we consider the same instances of (NEP) as above, but
consider the rst 200 coecients of each capacity constraint to be uncertain. This results
in signicantly more dicult instances for which exact solution or tight approximation are
too complex in most cases. Therefore, we use these instances to highlight the capability
of our framework to yield high quality solutions already for coarse approximations of the
robust counterpart. Table 12.6 shows the comparison of an approximation with  = 1, the
coarsest possible approximation, with an exact solution of each instance.

The results show a very clear picture. Whenever the approximate robust counterpart can
be solved to optimality, the solution time is much shorter than that for an exact solution
of the problem. Even when optimal solution of the approximation is not possible, the
solution obtained within the time limit is of high quality. Only for the robustied instance
Deutschland-700, the approximation did not yield a feasible solution within 1 h due to
the complexity of the root relaxation. Considering the adjusted measure of quality, the
obtained approximate solutions always stay within about 3 % of the optimal solution. This
even holds in cases where the exact model could only produce an upper bound on the prot
and no feasible solution. Altogether, we see that our framework is well suited to produce
robust solutions of high quality within short time.
188 12. Computational Assessment of Approximate Robust Optimization

LR-1 QR-GL
Solution time or Solution quality Solution time or
Instance Remaining gap plain / adjusted Remaining gap

SH-HA-12 8s 0.05 % / 1.5 % 25 s


B-BB-MV-5 12 s < 0.01 % / 0.23 % 64 s
NS-BR-180 0.16 % < 0.17 % / < 1.56 % 0.26 %
HE-RP-SL-295 62 s 0.03 % / 0.31 % 553 s
TH-S-SA-60 1465 s 0.06 % / 3.22 % 0.1 %
BA-BW-165 0.32 % < 0.36 % / < 2.64 % −
NRW-90 0.05 % < 0.06 % / < 1.03 % −
Deutschland-700 − − −

Table 12.6.: Performance of a coarse polyhedral approximation compared to Gurobi's gra-


dient linearization on railway instances with high uncertainty
Conclusions and Outlook

Three highly topical questions on the practical solution of network design problems have
been posed at the beginning of this thesis and have been the motivation for an extensive
study of algorithms to treat large-scale instances. Three dierent problem variants have
been subject to investigation, each of them with its own intrinsic challenges. And in all
three cases we have been able to nd answers allowing to make important steps towards
ecient solution strategies, spanning the whole range from exact over heuristic to approx-
imative approaches. These answers, together with the new research directions they lead
to, are summarized and evaluated in the following.

How can we devise ecient algorithms to solve multi-period network design problems?
 By decomposition along the timescale! Driven by a task set by our industry partner
GSV, we developed a compact multi-period planning model for the necessary expansion
of the German railway network until 2030. Our aim was to nd models and algorithms
to determine an optimal upgrade of the network and an optimal nancing plan to realize
it at the same time. This task is prototypical for a huge variety of investment planning
problems, not only limited to network problems, as the question of an optimal investment
schedule gains more and more interest in dierent application contexts. At the example of
Model (CFMNEP) for an optimal upgrade plan for the German railway network, we devel-
oped an algorithmic scheme to solve multi-period network design problems by a suitable
decomposition along the timescale. Our multiple-knapsack decomposition is a three-step
procedure: determination of an appropriate target network, evaluation of the chosen up-
grades with respect to their protability in each time period and a multiple-knapsack-type
subproblem to determine the nal upgrade schedule. In its basic version, this procedure is
an ecient heuristic which already leads to near-optimal solutions on the original problem
data provided by GSV. However, we were also able to show that a slight modication
allows to extended this method to an exact algorithm. It is an interesting new question
if this exact version of the method is practically applicable in broader problem contexts,
starting from general multi-period network design, but possibly also including other incre-
mental investment problems. Depending on the concrete problem, the heuristic scheme
might suer from a higher proportion of integer variables as well as integer variables in the
objective function, which it was not laid out for. The iterative character of its exact exten-
sion might then allow to correct initially imprecise evaluations of the available investments
very quickly. In any case, for the task of determining ecient network upgrades for railway
trac, our method proved to be a very useful planning tool. Summarizing the discussion
of its results in Chapter 5, we can say that our approach adresses the problem posed by
our partner GSV very well. This is not only our own opinion, but has been conrmed in
consultations with GSV itself, whose planners are very satised with the quality of the

© Springer Fachmedien Wiesbaden 2016


A. Bärmann, Solving Network Design Problems via Decomposition,
Aggregation and Approximation, DOI 10.1007/978-3-658-13913-1_14
190 Conclusions and Outlook

produced solutions. Most notably, our proposed expansion plan for Germany until 2030
provides a plausible answer to the increase in trac forecasted by GSV. This veries that
our model is an adequate representation of reality. Its compact reformulation and our
preprocessing routines showed to be important steps to enable its solvability. Moreover,
multiple-knapsack decomposition is demonstrated to be a fast method yielding high-quality
solutions for a problem which is dicult to solve via standard methods. All these features
together will allow GSV to benet from the software planning tool we developed on the
basis of these ndings. Thus, our work is a further example of the practical applicability
of discrete optimization for the solution of real-world planning problems.

How can we eciently deal with the large-scale networks arising in real-world prob-
lems?  By aggregation of the network structure! In the second part of this thesis, we
shifted our focus to a more general task. We were looking for an ecient way to cope with
the large-scale network structures present in many network design applications, including 
but by far not limited to  railway network design. Our idea was the derivation of a smaller
representation of the underlying graph by clustering the set of nodes, retaining as much
information as possible. The result was a new algorithmic framework for the solution of
network design problems based on iterative node aggregation. It relies on the interplay of a
master problem to propose network designs for the aggregated version of the network and
a subproblem to give a guideline for disaggregation in case of infeasibility. Based on this
principle, we rst developed an exact algorithm for network design without routing costs,
which is the easier case as the subproblem does not generate any costs. If integrated into
a branch-and-bound process, it takes the form of a cutting-plane procedure whose cutting
planes we could show to dominate Benders feasibility cuts. In a second step, we extended
this basic variant such that it supports routing costs on the arcs. This was possible via
the incorporation of lifted Benders optimality cuts to reformulate the routing costs within
the master problem. Again, we were able to derive an exact algorithm. In both cases, our
computational results conrm the superior performance of the aggregative approach over
a standard solver. It is considerably faster on a variety of problem instances of dierent
structure, including scale-free networks and railway network instances. We could show that
our method allows for a signicant reduction of the size of the networks under considera-
tion. But even when we had to disaggregate until reaching the original graph, our method
was mostly faster than solving the whole problem at once. The results thus indicate a
very benecial use of aggregation for network design problems. A natural direction for
future research is again to assess the applicability to more general problems. Aggregation
of the network structure might prove a powerful algorithmic framework for vehicle rout-
ing problems or scheduling problems in time-expanded networks as well. The transfer to
these problem contexts does not necessarily have to preserve the exactness of the method
to make its use benecial. Practical treatment of these very hard optimization problems
would also prot from ecient heuristics based on aggregation. Seen from a higher per-
spective, our research has paved to way to transferring the principle of iterative renement
from numerical mathematics, where it is omnipresent, to combinatorial optimization.

How can the robust counterpart of a network design problem be solved eciently in
the common case of ellipsoidal uncertainty?  By polyhedral approximation of the
Conclusions and Outlook 191

uncertainty sets! For the last part of the present work, we turned our attention to an
even more general problem setting. In railway network design, the inherent uncertainty in
the demand forecast may largely aect the eciency of the nal solution. Such uncertainty
is often modelled via ellipsoidal uncertainty sets as they allow to incorporate information
on the distribution of the uncertain coecients. However, the arising robust counterpart
is much harder to solve than in the case of polyhedral uncertainty. This was our motiva-
tion to develop a linear approximation of the second-oder cone robust counterpart via a
polyhedral approximation of the ellipsoidal uncertainty sets. The result is a general ap-
proximation framework that is not only applicable to railway network design, but to any
mixed-integer linear program under Ben-Tal-Nemirovski uncertainty. Employing the outer
approximation of the second-order cone developed by Ben-Tal and Nemirovski as well as
Glineur allows for a compact polyhedral approximation of the ellipsoidal uncertainty. The
same is possible via our new inner approximation of the second-order cone which is optimal
with respect to the size of the linear system and which possesses the same accuracy guar-
antees as the outer approximation. Our computational results show that this approach
is competitive to gradient-based approximations and to interior-point methods. In many
cases, it signicantly outperforms them. For the instances from railway network design
in particular, signicant savings in computation time are possible at very small losses in
accuracy only. Further computation time can be saved when using our approximation
as a heuristic by prescribing a lower approximation accuracy. The corresponding results
still have a very high quality. When considering constraints with many (> 100) uncertain
coecients, we noticed that the linear relaxations in the branch-and-bound tree become
very hard to solve in our framework. This leads to a very interesting idea to improve our
framework further by making use of our new inner approximation of the second-order cone.
While outer approximation of the uncertainty set entails a restriction of the original prob-
lem by making the robust counterpart slightly more conservative, the inner approximation
leads to a tight relaxation of the problem. This may be exploited to obtain a cutting-plane
procedure that starts with a coarse approximation of the uncertainty set and gradually
renes its representation by separating new scenarios in the process. Consequently, the
framework presented in this thesis is the rst step to a much more ecient treatment of
second-order cone robust optimization than possible before.

Network design problems have been of great interest to the optimization community in
recent decades. It is their high exibility for modelling application problems together
with their rich combinatorial structure that makes their study fascinating from both a
practical and a theoretical point of view. This work set out to answer a variety of trending
research questions in their algorithmic treatment, and its results lead to ecient solution
strategies in dierent problem contexts. At the same time, they bring up new compelling
directions for future research to build upon our ndings. Altogether, we are very condent
that network design will remain an important topic in combinatorial optimization in the
decades to come.
Bibliography

AEG (Fassung vom 07. August 2013). Allgemeines Eisenbahngesetz (AEG). [Link]
[Link]/bundesrecht/aeg\_1994/[Link].

Ahuja, R. K., Magnanti, T. L., and Orlin, J. B. (1993). Network Flows. Prentice Hall, Inc.

Albert, R. and Barabási, A.-L. (2002). Statistical mechanics of complex networks. Reviews
of Modern Physics, 74:4797.

ARRIVAL (2015). Homepage of Project ARRIVAL: Algorithms for Robust and online
Railway optimization: Improving the Validity and reliAbility of Large scale systems:
[Link]

Atamtürk, A. (2003). On capacitated network design cut-set polyhedra. Mathematical


Programming B, 92:425437.

Atamtürk, A. and Rajan, D. (2002). On splittable and unsplittable ow capacitated net-
work design arc-set polyhedra. Mathematical Programming A, 92:315333.

Balakrishnan, A. (1987). LP extreme points and cuts for the xed-charge network design
problem. Mathematical Programming, 39:263284.

Balakrishnan, A., Magnanti, T. L., and Mirchandani, P. (1997). Network design. In


Dell'Amico, M., Maoli, F., and Martello, S., editors, Annotated Bibliographies in Com-
binatorial Optimization, pages 311334. John Wiley & Sons, Inc.

Balakrishnan, A., Magnanti, T. L., and Wong, R. T. (1989). A dual-ascent procedure for
large-scale uncapacitated network design. Operations Research.

Balas, E. (1965). Solution of large-scale transportation problems through aggregation.


Operations Research, 13(1):8293.

Baxter, M., Elgindy, T., Ernst, A. T., Kalinowski, T., and Savelsbergh, M. W. P. (2014).
Incremental network design with shortest paths. European Journal of Operational Re-
search, 238:675684.

Beck, M. J. (2013). Kapazitätsmanagement und Netzentwicklung: Erfahrungen


mit Kompromissen zum Fahrplan 2018. Presentation by DB Netz AG, avail-
able [Link]
at:
com_docman&Itemid=64&task=doc_download&gid=84.

Ben-Tal, A., El Ghaoui, L., and Nemirovski, A. (2009). Robust Optimization. Princeton
University Press.

© Springer Fachmedien Wiesbaden 2016


A. Bärmann, Solving Network Design Problems via Decomposition,
Aggregation and Approximation, DOI 10.1007/978-3-658-13913-1
194 Bibliography

Ben-Tal, A. and Nemirovski, A. (2000). Robust solutions of linear programming problems


contaminated with uncertain data. Mathematical Programming A, 88:411424.

Ben-Tal, A. and Nemirovski, A. (2001). On polyhedral approximations of the second-order


cone. Mathematics of Operations Research, 26:193205.

Benders, J. F. (1962). Partitioning procedures for solving mixed-variables programming


problems. Numerische Mathematik, 4(1):238252.

Berndt, T. (2001). Eisenbahngüterverkehr. Teubner.

Bertsimas, D., Brown, D., and Caramanis, C. (2011). Theory and applications of robust
optimization. SIAM Review, 53(3):464501.

Bertsimas, D., Pachamanovab, D., and Sim, M. (2004). Robust linear optimization under
general norms. Operations Research Letters, 32:510516.

Bertsimas, D. and Sim, M. (2003). Robust discrete optimization and network ows. Math-
ematical Programming B, 98:4971.

Bertsimas, D. and Sim, M. (2004). The price of robustness. Operations Research, 52(1):35
53.

Bertsimas, D. and Sim, M. (2006). Tractable approximations to robust conic optimization


problems. Mathematical Programming B, 107:536.

Bienstock, D. and Günlük, O. (1996). Capacitated network design  polyhedral structure


and computation. INFORMS Journal on Computing, 8(3):243259.

Bienstock, D., Raskina, O., Saniee, I., and Wang, Q. (2006). Combined network design
and multiperiod pricing: Modeling, solution techniques, and computation. Operations
Research, 54(2):261276.

Blanco, V., Puerto, J., and Ramos, A. B. (2011). Expanding the Spanish high-speed
railway network. Omega, 39:138150.

Boland, N., Ernst, A., Kalinowski, T., Rocha de Paula, M., Savelsbergh, M., and Singh,
G. (2013). Time aggregation for network design to meet time-constrained demand. In
Piantadosi, J., Anderssen, R. S., and Boland, J., editors, 20th International Congress
on Modelling and Simulation (MODSIM2013), pages 32813287.

Bonnans, J. F., Gilbert, J. C., Lemaréchal, C., and Sagastizábal, C. A. (2009). Numerical
Optimization  Theoretical and Practical Aspects. Springer.

Borndörfer, R., Erol, B., Graagnino, T., Schlechte, T., and Swarat, E. (2014a). Optimiz-
ing the simplon railway corridor. Annals of Operations Research, 218(1):93106.

Borndörfer, R., Reuther, M., and Schlechte, T. (2014b). A coarse-to-ne approach to the
14th Workshop on Algorithmic Approaches for
railway rolling stock rotation problem. In
Transportation Modelling, Optimization, and Systems (ATMOS'14), volume 42, pages
7991.
Bibliography 195

Breimeier, R. and Konanz, W. (1994). Die wirtschaftliche Zugführung als Instru-


ment der langfristigen Infrastrukturplanung. Schriftenreihe der Deutschen Verkehrswis-
senschaftlichen Gesellschaft e.V. (DVWG), 178:168185.

BSchWAG (Fassung vom 31. Oktober 2006). Bundesschienenwegeausbaugesetz


(BSchWAG). [Link]
pdf.

BVWP (Fassung vom 02. Juli 2003). Bundesverkehrswegeplan 2003. http:


//[Link]/SharedDocs/DE/Anlage/VerkehrUndMobilitaet/Schiene/2003/
bundesverkehrswege-plan-2003-beschluss-der-bundesregierung-vom-02-juli-
[Link]?__blob=publicationFile.

Bärmann, A., Heidt, A., Martin, A., Pokutta, S., and Thurner, C. (2015a). Polyhedral ap-
proximation of ellipsoidal uncertainty sets via extended formulations: A computational
case study. to appear in Computational Management Science.

Bärmann, A., Liers, F., Martin, A., Merkert, M., Thurner, C., and Weninger, D. (2015b).
Solving network design problems via iterative aggregation. Mathematical Programming
Computation, 7(2):189217.

Caimi, G. C. (2009). Algorithmic Decision Support for Train Scheduling in a Large and
Highly Utilised Railway Network. PhD thesis, Eidgenössische Technische Hochschule
Zürich.

Cao, M., Wang, X., Kim, S.-J., and Madihian, M. (2007). Multi-hop wireless backhaul
networks: a cross-layer design paradigm. IEEE Journal on Selected Areas in Communi-
cations, 25(4):738748.

Chang, S.-G. and Gavish, B. (1993). Telecommunications network topological design and
capacity expansion: Formulations and algorithms. Telecommunication Systems, 1(1):99
131.

Chang, S.-G. and Gavish, B. (1995). Lower bounding procedures for multiperiod telecom-
munications network expansion problems. Operations Research, 43(1):4357.

Chopra, S., Gilboa, I., and Sastry, S. T. (1998). Source sink ows with capacity installation
in batches. Discrete Applied Mathematics, 85(3):165192.

Chouman, M., Crainic, T. G., and Gendron, B. (2009). A cutting-plane algorithm based on
cutset inequalities for multicommodity capacitated xed-charge network design. Tech-
nical Report CIRRELT-2009-20, Centre de recherche sur les transports, Université de
Montréal.

Christodes, N. and Brooker, P. (1974). Optimal expansion of an existing network. Math-


ematical Programming, 6:197211.

Chvátal, V. and Hammer, P. L. (1977). Aggregation of inequalities in integer programming.


In Hammer, P., Johnson, E., Korte, B., and Nemhauser, G., editors, Studies in Integer
Programming, volume 1 of Annals of Discrete Mathematics, pages 145162. Elsevier.
196 Bibliography

Conforti, M., Cornuéjols, G., and Zambelli, G. (2010). Extended formulations in combi-
natorial optimization. 4OR, 8(1):148.

Costa, A. M. (2005). A survey on benders decomposition applied to xed-charge network


design problems. Computers and Operations Research, 32(6):14291450.

Costa, A. M., Cordeau, J.-F., and Gendron, B. (2009). Benders, metric and cutset inequal-
ities for multicommodity capacitated network design. Computational Optimization and
Applications, 42:371392.

Crainic, T. G., Li, Y., and Toulouse, M. (2006). A rst multilevel cooperative algorithm
for capacitated multicommodity network design. Computers & Operations Research,
33(9):26022622.

Dantzig, G. B. and Wolfe, P. (1960). Decomposition principle for linear programs. Opera-
tions Research, 8(1):101111.

DB Netz AG (2013). Netzkonzeption 2030: Zielnetz der DB Netz AG für die Schienenin-
frastruktur im Jahr 2030. Information Booklet.

Demetrescu, C., Goldberg, A., and Johnson, D. (2006). 9th DIMACS implementation
challenge  shortest paths. [Link]
Dempster, M. A. H. and Thompson, R. T. (1998). Parallelization and aggregation of nested
Benders decomposition. Annals of Operations Research, 81:163187.

Desrosiers, J. and Lübbecke, M. E. (2005). A primer in column generation. In Desaulniers,


G., Desrosiers, J., and Solomon, M., editors, Column Generation, pages 132. Springer
US.

Dezs®, B., Jüttner, A., and Kovácsa, P. (2011). Lemon  an open source c++ graph
template library. Electronic Notes in Theoretical Computer Science, 264(5):2345.

Dogan, K. and Goetschalckx, M. (1999). A primal decomposition method for the integrated
design of multi-period production-distribution systems. IIE Transactions, 31(11):1027
1036.

Dolan, E. and Moré, J. (2002). Benchmarking optimization software with performance


proles. Mathematical Programming A, 91(2):201213.

Dudkin, L., Rabinovich, I., and Vakhutinsky, I. (1987). Iterative aggregation theory, volume
111 of Pure and Applied Mathematics. Dekker.

Dutta, A. and Lim, J.-I. (1992). A multi-period capacity planning model for backbone
computer communication networks. Operations Research, 40(4):689705.

EBO (Fassung vom 25. Juli 2012). Eisenbahn-Bau- und Betriebsordnung (EBO). http:
//[Link]/bundesrecht/ebo/[Link].
Engel, K., Kalinowski, T., and Savelsbergh, M. W. P. (2013). Incremental network design
with minimum spanning trees. Technical report, Universität Rostock, Germany, and
University of Newcastle, Australia.
Bibliography 197

Fiorini, S., Rothvoÿ, T., and Tiwary, H. (2012). Extended formulations for polygons.
Discrete Computational Geometry, 48(3):658668.

Fischetti, M. and Monaco, M. (2009). Light robustness. In Ahuja, R. K., Möhring, R. H.,
Robust and Online Large-Scale Optimization, volume 5868
and Zaroliagis, C. D., editors,
of Lecture Notes in Computer Science, pages 6184. Springer.
Francis, V. E. (1985). Aggregation of Network Flow Problems. PhD thesis, University of
California.

Frangioni, A. (2005). About lagrangian methods in integer optimization. Annals of Oper-


ations Research, 139:163193.

Frangioni, A. and Gendron, B. (2009). 0-1 reformulations of the multicommodity capaci-


tated network design problem. Discrete Applied Mathematics, 157:12291241.

Frangioni, A. and Gendron, B. (2013). A stabilized structured Dantzig-Wolfe decomposi-


tion method. Mathematical Programming B, 140:4576.

Gamst, M. and Spoorendonk, S. (2013). An exact approach for aggregated formulations.


Technical Report DTU Management Engineering Report 3.2013, Technical University of
Denmark.

Garcia, B.-L., Mahey, P., and LeBlanc, L. J. (1998). Iterative improvement methods
for a multiperiod network design problem. European Journal of Operational Research,
110:150165.

Garey, M. R. and Johnson, D. S. (1979). Computers and Intractability: A Guide to the


Theory of NP-Completeness. W.H. Freeman and Company, New York.

Gay, M. (1985). Electronic mail distribution of linear programming test problems. Math-
ematical Programming Society COAL Bulletin, 13:1012. Data available at http:
//[Link]/netlib/lp.
Geisberger, R., Sanders, P., Schultes, D., and Vetter, C. (2012). Exact routing in large
road networks using contraction hierarchies. Transportation Science, 46(3):388404.

Gendreau, M., Potvin, J.-Y., Smires, A., and Soriano, P. (2006). Multi-period capacity ex-
pansion for a local access telecommunications network. European Journal of Operational
Research, 172(3):10511066.

Gendron, B. (1994). Nouvelles méthodes de résolution de problèmes de conception de


réseaux et leur implantation en environnement parallèle. PhD thesis, Université de Mon-
tréal.

Gendron, B., Crainic, T. G., and Frangioni, A. (1999). Multicommodity capacitated net-
work design. In Sansò, B. and Soriano, P., editors, Telecommunications Network Plan-
ning, pages 119. Springer.

Gendron, B., Crainic, T. G., and Frangioni, A. (2001). Bundle-based relaxation methods for
multicommodity capacitated xed charge network design. Discrete Applied Mathematics,
112:7399.
198 Bibliography

Gendron, B. and Larose, M. (2014). Branch-and-price-and-cut for large-scale multicom-


modity capacitated xed-charge network design. EURO Journal on Computational Op-
timization, 2(1-2):5575.

Georion, A. M. (1972). Generalized Benders decomposition. Journal of Optimization


Theory and Applications, 10(4):237260.

Georion, A. M. (1974). Lagrangean relaxation for integer programming. Mathematical


Programming Study, 2:82114.

Glineur, F. (2000). Computational experiments with a linear approximation of second-


order cone optimization. Image Technical Report 001, Faculté Polytechnique de Mons.

Glineur, F. (2001). Conic optimization: an elegenat framework for convex optimization.


Belgian Journal of Operations Research, Statistics and Computer Science, 41(1-2):528.

Gupta, A. and Könemann, J. (2011). Approximation algorithms for network design: A


survey. Surveys in Operations Research and Management Science, 16(1):320.

Gurobi Optimization, Inc. (2014). Gurobi optimizer reference manual. [Link]


[Link].
Hallefjord, A. and Storoy, S. (1990). Aggregation and disaggregation in integer program-
ming problems. Operations Research, 38(4):619623.

Hellstrand, J., Larsson, T., and Migdalas, A. (1992). A characterization of the uncapaci-
tated network design polytope. Operations Research Letters, 12(3):159163.

Hewitt, M., Nemhauser, G. L., and Savelsbergh, M. W. P. (2010). Combining exact and
heuristic approaches for the capacitated xed-charge network ow problem. INFORMS
Journal on Computing, 22(2):314325.

Holmberg, K. and Hellstrand, J. (1998). Solving the uncapacitated network design problem
by a lagrangean heuristic and branch-and-bound. Operations Research, 46:2.

Holmberg, K. and Yuan, D. (1998). A lagrangean approach to network design problems.


International Transactions in Operational Research, 5(6):529539.

Holmberg, K. and Yuan, D. (2003). A multicommodity network-ow problem with side


constraints on paths solved by column generation. INFORMS Journal on Computing,
15(1):4257.

Hörl, B. (1998). Engpassbeseitigende Investitionsmaÿnahmen auf Schienenstrecken und


deren Bewertung, volume 3 of IVS-Schriften. Österreichischer Kunst- und Kulturverlag.
Kaibel, V. (2011). Extended formulations in combinatorial optimization. Optima, 85:27.

Kaibel, V. and Pashkovich, K. (2011). Constructing extended formulations from reec-


tion relations. In Günlük, O. and Woeginger, G. J., editors, Integer Programming and
Combinatoral Optimization, volume 6655 of Lecture Notes in Computer Science, pages
287300. Springer.

Kalinowski, T., Matsypurab, D., and Savelsbergh, M. W. (2015). Incremental network


design with maximum ows. European Journal of Operational Research, 242:5162.
Bibliography 199

Karwan, M. H. and Rardin, R. L. (1979). Some relationships between Lagrangian and


surrogate duality in integer programming. Mathematical Programming, 17(1):320334.

Kerivin, H. and Mahjoub, A. R. (2005). Design of survivable networks: A survey. Networks,


46(1):121.

Kettner, M. (2005). Netz-Evaluation und Engpassbehandlung mit makroskopischen Mod-


ellen des Eisenbahnbetriebs. PhD thesis, Universität Hannover.

Kim, B. J., Kim, W., and Song, B. H. (2008). Sequencing and scheduling highway network
expansion using a discrete network design model. The Annals of Regional Science,
42(3):621642.

Kim, D., Barnhart, C., Ware, K., and Reinhardt, G. (1999). Multimodal express package
delivery: A service network design application. Transportation Science, 33(4):391407.

Kouassi, R., Gendreau, M., Potvin, J.-Y., and Soriano, P. (2009). Heuristics for multi-
period capacity expansion in local telecommunications networks. Journal of Heuristics,
15:381402.

Kubat, P. and Smith, J. M. (2001). A multi-period network design problem for cellular
telecommunication systems. European Journal of Operational Research, 134(2):439456.

Kuby, M., Xu, Z., and Xie, X. (2001). Railway network design with multiple project stages
and time sequencing. Journal of Geographical Systems, 3:2547.

Lai, Y.-C. and Shih, M.-C. (2013). A stochastic multi-period investment selection model
to optimize strategic railway capacity planning. Journal of Advanced Transportation,
47:281296.

Lardeux, B., Nace, D., and Geard, J. (2007). Multiperiod network design with incremental
routing. Networks, 50(1):109117.

Lee, S. (1975). Surrogate Programming by Aggregation. PhD thesis, University of California.

Iterative Aggregation und Mehrstuge Entscheidungsmodelle: Einord-


Leisten, R. (1995).
nung in den Planerischen Kontext, Analyse anhand der Modelle der Linearen Program-
mierung und Darstellung am Anwendungsbeispiel der Hierarchischen Produktionspla-
nung. Produktion und Logistik. Physica-Verlag.

Leisten, R. (1998). An LP-aggregation view on aggregation in multi-level production


planning. Annals of Operations Research, 82:413434.

Lemaréchal, C. (2001). Lagrangian relaxation. In Jünger, M. and Naddef, D., editors,


Computational Combinatorial Optimization, volume 2241 of Lecture Notes in Computer
Science, pages 112156. Springer.

Liebchen, C., Lübbecke, M., Möhring, R., and Stiller, S. (2009). The concept of recov-
erable robustness, linear programming recovery, and railway applications. In Ahuja,
R. K., Möhring, R. H., and Zaroliagis, C. D., editors, Robust and Online Large-Scale
Optimization, volume 5868 of Lecture Notes in Computer Science, pages 127. Springer.
200 Bibliography

Linderoth, J., Margot, F., and Thain, G. (2009). Improving bounds on the football pool
problem by integer programming and high-throughput computing. INFORMS Journal
on Computing, 21(3):445457.

Litvinchev, I. and Tsurkov, V. (2003). Aggregation in Large-Scale Optimization. Applied


Optimization. Springer.

Ljubi¢, I., Putz, P., and Salazar-González, J.-J. (2012). Exact approaches to the single-
source network loading problem. Network, 59(1):89106.

Lumbreras, S. and Ramos, A. (2013). Optimal design of the electrical layout of an oshore
wind farm applying decomposition strategies. IEEE Transactions on Power Systems,
28:11341441.

Macedo, R., Alves, C., de Carvalho, J. V., Clautiaux, F., and Hana, S. (2011). Solving
the vehicle routing problem with time windows and multiple routes exactly using a
pseudo-polynomial model. European Journal of Operational Research, 214(3):536545.

Magnanti, T. L. and Mirchandani, P. (1993). Shortest paths, single origin-destination


network design, and associated polyhedra. Networks, 23(2):103121.

Magnanti, T. L., Mireault, P., and Wong, R. T. (1986). Tailoring Benders decomposition
for uncapacitated network design. In Gallo, G. and Sandi, C., editors, Netow at Pisa,
volume 26 of Mathematical Programming Studies, pages 112154. Springer.

Magnanti, T. L. and Wong, R. T. (1984). Network design and transportation planning:


Models and algorithms. Transportation Science, 18(1):155.

Marín, Á. and Jaramillo, P. (2008). Urban rapid transit network capacity expansion.
European Journal of Operational Research, 191:4560.

McDaniel, D. and Devine, M. (1977). A modied Benders' partitioning algorithm for mixed
integer programming. Management Science, 24(3):312319.

Melo, M. T., Saldanha da Gama, F., and Silva, M. M. (2005). A primal decomposition
scheme for a dynamic capacitated phase-in/phase-out location problem. Technical Re-
port CIO  Working paper 9/2005, Centro de Investigação Operacional, University of
Lisbon.

Minoux (1989). Network synthesis and optimum network design problems: Models, solu-
tion methods and applications. Networks, 19(3):313360.

Minoux, M. (1987). Network synthesis and dynamic network optimization. Annals of


Discrete Mathematics, 31:283324.

Newman, A. M. and Kuchta, M. (2007). Using aggregation to optimize long-term pro-


duction planning at an underground mine. European Journal of Operational Research,
176(2):12051218.

Orlowski, S., Pióro, M., Tomaszewski, A., and Wessäly, R. (2010). SNDlib 1.0Survivable
Network Design Library. Networks, 55(3):169297.
Bibliography 201

Papadimitriou, D. and Fortz, B. (2014). Time-dependent combined network design and


routing optimization. In IEEE International Conference on Communications (ICC)
2014, pages 11241130.

Petersen, E. R. and Taylor, A. J. (2001). An investment planning model for a new north-
central railway in Brazil. Transportation Research Part A: Policy and Practice, 35:847
862.

Pickavet, M. and Demeester, P. (1999). Long-term planning of WDM networks: A com-


parison between single-period and multi-period techniques. Photonic Network Commu-
nications, 1(4):331346.

Pióro, M. and Medhi, D. (2004). Routing, Flow, and Capacity Design in Communica-
tion and Computer Networks. The Morgan Kaufmann Series in Networking. Morgan
Kaufmann, Inc.

Poss, M. and Raack, C. (2013). Ane recourse for the robust network design problem:
Between static and dynamic routing. Networks, 61(2):180198.

Pruckner, M., Thurner, C., Martin, A., and German, R. (2014). A coupled optimization
and simulation model for the energy transition in bavaria. In Fischbach, K., Groÿmann,
M., Krieger, U. R., and Staake, T., editors, Proceedings of the International Workshops
SOCNET 2014 and FGENET 2014, pages 97104.

Raack, C. (2012). Capacitated Network Design  Multi-Commodity Flow Formulations,


Cutting Planes, and Demand Uncertainty. PhD thesis, Technische Universität Berlin.

Repolho, H. M., Antunes, A. P., and Church, R. L. (2013). Optimal location of railway
stations: The Lisbon-Porto high-speed rail line. Transportation Science, 47(3):330343.

Ridley, T. M. (1968). An investment policy to reduce the travel time in a transportation


network. Transportation Research, 2:409424.

Rosenberg, I. (1974). Aggregation of equations in integer programming. Discrete Mathe-


matics, 10(2):325341.

Ross, S. (2001). Strategische Infrastrukturplanung im Schienenverkehr. PhD thesis, Uni-


versität Köln.

Salvagnin, D. and Walsh, T. (2012). A hybrid MIP/CP approach for multi-activity shift
scheduling. In Milano, M., editor, Principles and Practice of Constraint Programming,
volume 7514 of Lecture Notes in Computer Science, pages 633646. Springer.

Schlechte, T., Borndörfer, R., Erol, B., Graagnino, T., and Swarat, E. (2011). Micro-
macro transformation of railway networks. Journal of Rail Transport Planning & Man-
agement, 1(1):3848.

Schriftenreihe der Deutschen


Schwanhäuÿer, W. (1995). Leistungsfähigkeit und Kapazität.
Verkehrswissenschaftlichen Gesellschaft e.V., DVWG Reihe B: Seminar, 178:317.
202 Bibliography

Schwanhäuÿer, W. (1997). German introductory report (untitled). In The European Con-


ference of Ministers of Transport (ECMT), Economic Research Centre: The Separation
of Operations from Infrastructure in the Provision of Railway Services, Report of the
103rd Round Table on Transport Economics, pages 552.

Scott, A. J. (1969). The optimal network problem: Some computational procedures. Trans-
portation Research, 3(2):201210.

Sellmann, M., Kliewer, G., and Koberstein, A. (2002). Lagrangian cardinality cuts and
variable xing for capacitated network design. In Möhring, R. and Raman, R., editors,
Algorithms  ESA 2002, volume 2461 of Lecture Notes in Computer Science, pages 845
858. Springer.

Sewcyk, B. (2004). Makroskopische Abbildung des Eisenbahnbetriebs in Modellen zur


langfristigen Infrastrukturplanung. PhD thesis, Universität Hannover.

Shapiro, J. F. (1984). A note on node aggregation and Benders' decomposition. Mathe-


matical Programming, 29:113119.

Shulman, A. and Vachani, R. (1993). An algorithm for capacity expansion of local access
networks. IEEE Transactions on Communications, 41(7):10631073.

Sivaraman, R. (2007). Capacity Expansion in Contemporary Telecommunication Networks.


PhD thesis, Massachusetts Institute of Technology.

Spönemann, J. C. (2013). Network Design for Railway Infrastructure by means of Linear


Programming. PhD thesis, Rheinisch-Westfälische Technische Hochschule Aachen.

Stairs, S. (1968). Selecting an optimal trac network. Journal of Transport Economics


and Policy, 2(2):218231.

Stallaert, J. (2000). Valid inequalities and separation for capacitated xed charge ow
problems. Discrete Applied Mathematics, 98:265274.

Statistisches Bundesamt (2014). Homepage of the Federal Statistical Oce (Statistisches


Bundesamt): [Link]

Thapalia, B. K., Wallace, S. W., Kaut, M., and Crainic, T. G. (2012). Single source single-
commodity stochastic network design. Computational Management Science, 9:139160.

Toriello, A., Nemhauser, G., and Savelsbergh, M. (2010). Decomposing inventory routing
problems with approximate value functions. Naval Research Logistics, 57(8):718727.

Trukhanov, S., Ntaimo, L., and Schaefer, A. (2010). Adaptive multicut aggregation for
two-stage stochastic linear programs with recourse. European Journal of Operations
Research, 206:395406.

Umweltbundesamt (2009). Strategie für einen nachhaltigen Güterverkehr. available


at: [Link]
long/[Link].
Bibliography 203

Umweltbundesamt (2010). Schienennetz 2025/2030  Ausbaukonzeption für einen


leistungsfähigen Schienengüterverkehr in Deutschland. [Link]
available at:
[Link]/sites/default/files/medien/461/publikationen/[Link].
van Hoesel, S. P. M., Koster, A. M. C. A., van de Leensel, R. L. M. J., and Savelsbergh,
M. W. P. (2004). Bidirected and unidirected capacity installation in telecommunication
networks. Discrete Applied Mathematics, 133:103121.

Vielma, J., Ahmed, S., and Nemhauser, G. L. (2008). A lifted linear programming branch-
and-bound algorithm for mixed integer conic quadratic programs. INFORMS Journal
on Computing, 20(3):438450.

Vieregg, M. (1995). Ezienzsteigerung im Schienenpersonenfernverkehr. PhD thesis,


Ludwigs-Maximilians-Universität München.

Zipkin, P. H. (1977). Aggregation in Linear Programming. PhD thesis, Yale University.

Das könnte Ihnen auch gefallen