Research
Papers and benchmarks about how people move around Europe, written from the data published at /data and released under the same licence. Everything here states its version, its date and what it rests on.
Papers
Published, with the full text on this site as well as a PDF.
Shramish Kafle, Orlero
Abstract
Journey planners and timetable feeds report in-vehicle time. Travellers experience door-to-door time, which adds access, egress and the waiting that a filed connection imposes. This paper asks how far the two diverge across European corridors, and whether the divergence is large enough to change which mode is ranked fastest. Using a snapshot of 936 filed itineraries on 493 European corridors, every itinerary is decomposed into in-vehicle time, filed interchange time and a parameterised access and egress allowance. In-vehicle time accounts for a median of 77% of door-to-door time, with an interquartile range of 57 to 87.3%. On the 79 corridors where two or more modes are filed, the mode ranked fastest changes on 13 of them, 16.5%, when the clock changes from in-vehicle to door to door. The reversals are produced almost entirely by filed interchange time rather than by the access and egress allowance: every one of the 13 reversals also occurs when the allowance is set to zero, and the reversal rate moves only between 16.5 and 20.3% across the full range of allowances tested. A planner can therefore correct most of the bias in in-vehicle time using data it already holds, without first resolving the access-time parameter that the literature disagrees about.
In progress
Work that exists as code and data and has produced no result yet.
Grounded trip planning: a hallucination benchmark
No model has been scored yet.
It asks whether an assistant invents departures. The set is built so that saying "I do not know" is the right answer to a good share of it: real ports with no water route between them, routes that ran nothing on the day asked about, a carrier named on a corridor it does not serve, a place name that means two places.
- 398questions
- 10strata
- 100in the development split
- 298held for test
398 questions across 10 strata, drawn deterministically from one seed and split 100 for development and 298 for test. Ground truth comes from filed timetables, operator-published sailings and a carrier graph, with a scoring rubric written to be reproducible between two readers.
The run is the next step.
How this is published
- Attribution, on the paper and on the data under it
- The text, the tables, the figures and the released data are all CC BY 4.0. Quote it, redraw it, reuse it commercially. Name Orlero and link back.
- The data is released before the claim is made
- Every paper here names the datasets it read, each of which is published at /data with a version and a checksum. The analysis reads those files and writes both the tables in the PDF and the figures on the page, so the two renditions cannot state different numbers.
- The numbers are generated, not transcribed
- The page and the PDF are built from one results file that an analysis script writes from the committed data, and a test recomputes the whole analysis and fails if either has drifted from it.
- No DOI yet
- Zenodo deposits are prepared and unpublished, so cite the URL and the version.
- The competing interest is stated on the paper
- The author is the founder of Orlero, whose ranking method one of these papers is partly a critique of.