Open data

Open datasets

Four datasets about how people move around Europe, published as files rather than as charts. Every row says whether it was measured, filed, observed or derived.

Read the evidence column first

Every row in every one of these files carries an evidence value, and it decides what the row can be used for. The four datasets do not sit at the same level.

Measured
A departure, a duration or a price as the operator published it on the day it was read. Nothing was worked out.
Filed
What an operator says a service will do. A filed departure is a plan, not an observation.
Observed
A connection some source saw run, recorded in a feed or a public register. Weaker than measured, because the record is second hand, and stronger than derived, because something ran.
Derived
Worked out by us from other values: a distance from two sets of coordinates, a carrier assumed to serve a corridor because its published service area covers both ends, a gram count from a published factor.

A derived value stays derived in the file, in the page above it and in the structured data.

The four datasets

Ordered by the strength of what they carry rather than by size.

  1. Which boats sail between which ports, when, run by whom, and what the operator charged.

    Every one of the 2,709 sailings is measured: the departure, the arrival, the vessel and the fare as 120 operators published them, read on four probed dates and written down unaltered. Fares are assumed to be in euro.

    • 2,709sailings
    • 2,190crossings
    • 120operators
    • 65countries

    Version 1.0.0, rebuilt 12 September 2026, 6,698 rowsFormats: CSV, JSON

    Open the dataset
  2. What the timetable says a city to city journey looks like, leg by leg, with calling points and platforms.

    All 3,025 legs are filed: published plans, not observations of services that ran. 545 of 1,254 corridors came back with anything at all.

    • 545corridors
    • 1,035itineraries
    • 3,025legs
    • 184operators

    Version 1.0.0, rebuilt 16 September 2026, 19,195 rowsFormats: ZIP, GTFS, CSV

    Open the dataset
  3. Which carrier serves which pair of places, across every mode, worldwide.

    164,657 of 293,523 legs are mapped from an operator’s published service area covering both ends, which is evidence the carrier serves the corridor. The remaining 128,866 are read from a feed or a register.

    • 293,523legs
    • 2,093operators
    • 11,241places
    • 211countries

    Version 1.0.0, rebuilt 14 September 2026, 307,072 rowsFormats: CSV, PARQUET, JSON

    Open the dataset
  4. How many grams of CO2e a passenger kilometre carries, by mode, and where each factor came from.

    Each figure is a published factor multiplied by a modelled distance. 8 of the 8 rows in the factor table reproduce from a source we can name, and the table prints that status in the row.

    • 6factors in use
    • 8,994corridors with more than one mode
    • 18,537corridor and mode rows
    • 114,524corridors in the network

    Version 2.0.0, rebuilt 18 September 2026, 18,545 rowsFormats: CSV, JSON

    Open the dataset

True of all four

The parts that do not change between one dataset and the next.

Attribution, and that is the whole licence
Every file is CC BY 4.0. Use it commercially, redraw it, build on it, sell what you build. Name Orlero and link back.
A Frictionless Data Package, not a loose CSV
Each package carries a datapackage.json with a Table Schema per file, so a reader knows the type and the unit of every column before opening it. Dates are ISO 8601 with an offset, countries are ISO 3166-1 alpha-2, money is in integer minor units with its ISO 4217 code, and coordinates are WGS84.
Four provenance columns on every row
source, source_url, fetched_at and evidence, in every table in all four packages. A row lifted out of context still says where it came from, when it was read and how strong the claim is.
Checksums and a changelog
Every package ships checksums.txt with a SHA-256 per file, a README, a CHANGELOG and a semantic version. A rebuild that changes a column meaning moves the major version.
Missing data is an empty cell
A blank means the value is not held, and a zero means zero. Never a zero in place of a blank, and never the letters N slash A.

https://creativecommons.org/licenses/by/4.0/

How to cite

Each dataset page carries its own formatted citation and a BibTeX entry to copy. They follow one shape, which for the ferry package reads:

Orlero (2026). Scheduled ferry crossings with operator-published fares (Version 1.0.0) [Data set]. https://orlero.com/data/ferry-crossings

No DOI yet. Cite the URL above.

What is not here

The gaps worth knowing about before you build on any of this.

  • Every package is a dated snapshot rather than a live feed, and the date is printed on the dataset page, in the file and in the structured data.
  • Coverage is uneven by design: the ferry snapshot is dense around the Mediterranean, the timetables are European, and the operator graph reaches furthest and is checked least.
  • Fares are what an operator published on one day at one point in the booking window. They are not a fare curve and they are not an average.
  • The four packages are built by four separate scripts from overlapping sources, so a place can appear in two of them with two identifiers.
Plan with Ori

Ori, the trip planner

Ask it something like this