Uncategorized

How well does an algorithm know the way to school?

By Hyesop Shin, University of Auckland

Walking to school is one of the simplest ways for a child to incorporate physical activity into their daily routine. Geographers and health practitioners have long been interested in knowing exactly which streets children use because the route, not just the distance, determines what a child is exposed to along the way, from traffic and pollution to busy junctions that parents worry about.

But wait, how do we monitor that? To date, GPS tracking has been the most reliable way to trace children’s journeys because it is recorded on a 10 second epoch. That means, once somebody wears this device, it returns the geolocations every 10 seconds. However, these studies are expensive, ethically demanding (need consent from parents), have a high percentage of response errors, and are usually limited to small samples to represent a whole population.

Recent developments in routing engines such as Google Maps can overcome this limitation by predicting a plausible route home to school in seconds at almost no cost. The average routing engine generates the route that an algorithm would recommend to a reasonable traveller, not the route that a specific ten-year-old child actually walked. Children’s journeys are influenced by factors such as friends, habit, a desire to avoid busy junctions and a preference for fewer turns. None of these things can be taken into account in a shortest-path calculation. So the question for anyone using these tools is easy to state but hard to answer: How well do routing engines capture the reality of walking to school?

Putting the engines to the test

In a study published in the International Journal of Health Geographics, my colleagues and I from the Universities of Glasgow and Auckland put that question to the test. We compared the routes suggested by three widely available engines, Google, Mapbox, and Open Source Routing Machine (OSRM), against the GPS footprints of 233 children aged 10 to 11 years old walking to school across Scotland. For each child, we computed the overlap accuracy (OA), defined as the fraction of the actual GPS path that was reconstructed by the predicted route.

Figure 1 An example of the Overlapping Accuracy (OA) of three routing engines against the GPS trajectory (with black points) in a school journey, author provided.

On average, these engines replicated 62-67% of each child’s route, compared to about 45% for the traditional shortest-path model that GIScientists have relied on for decades. The improvement is significant because it shows commercial routing engines have a better understanding of how people move than a simple measurement of network distance. But still about a third of the journey was unseen.

Another finding was that the OA was not patterned by social or physical situation. There was no systematic difference by the neighbourhood deprivation index (Scottish Index of Multiple Deprivation), whether the home is located in a city or rural area, or the sex of the child. The engines were not quietly biased against poorer or denser neighbourhoods, which is reassuring for anyone concerned that an algorithm might serve some children better than others. Rather, the gap between prediction and reality seems to be a basic feature of how routing software models pedestrian movement, not a matter of who or where the child is. Rather, the messiness was more individual than structural. When we looked at the results for children alone, only about a third walked routes that were matched by all three engines.

Figure 2 Comparison of the OA of three routing engines in relation to real-world GPS data: overall (A), by sex (B), by socioeconomic (C), and by urban and non-urban (rural) settings (D), author provided.

What this means for studying movement

Geographers know this lesson well, in a new environment. The difference between a model and the world it represents is not something to apologise for, but rather a feature that can be measured and understood. Routing engines are good enough to be useful in practice. They pave the way for studying children’s active travel at a population scale. They could underpin larger and more representative studies of how children move through their neighbourhoods, used as a complement to GPS, not as a replacement for it.

The caution is equally clear. These tools should not be mistaken for ground truth, especially where a single decision rests on a single predicted route. In many places, including the US or New Zealand, school transport eligibility is determined by routing software that calculates the distance a child would have to travel, and even small differences between a modelled and an actual path can change who counts as living close enough to walk. When two apps disagree, knowing that each captures only about two-thirds of a real journey should make us slower to treat either as the final word.

Another point, which is very interesting, is that children did not necessarily leave from the same address each day. Therefore, each trajectory had to be matched with its submitted home address as its origin before any comparison could be made. They could go to school, their friends’ houses, their grandparents’ houses, or the houses of their aunts and uncles. It turns out that lived mobility is much messier than the tidy origin-to-destination logic built into the software.

The broader message is a quietly hopeful one for open and reproducible geography. The most accurate engine was not always the commercial one, and a free, open-source tool performed competitively throughout. Good research into how children move does not have to depend on proprietary platforms or costly fieldwork. It can be built, increasingly, from open data and open tools, provided we stay honest about the third of the journey we cannot yet see.


About the author: Hyesop Shin is a Lecturer in GIScience in the School of Environment at the University of Auckland. His research covers children’s active travel, transport equity and geospatial methods.

Suggested further reading

Harrison, F, Burgoine, T, Corder, K, van Sluijs, EMF and Jones, A (2014) How well do modelled routes to school record the environments children are exposed to?: a cross-sectional comparison of GIS-modelled and GPS-measured routes to school. International Journal of Health Geographics. Available from: https://doi.org/10.1186/1476-072X-13-5

Hasanzadeh et al (2022), Children’s physical activity and active travel: a cross-sectional study of activity spaces, sociodemographic and neighborhood associations. Children’s Geographies. Available from:  https://doi.org/10.1080/14733285.2022.2039901

Mavrogeni, M., van Dijk, J. & Longley, P. (2025) Understanding place-to-place interactions using flow patterns derived from in-app mobile phone location data. The Geographical Journal. Available from: https://doi.org/10.1111/geoj.70033

Panter, JR, Jones, AP, Van Sluijs, EMF and Griffin, SJ (2010) Neighborhood, route, and school environments and children’s active commuting. American Journal of Preventive Medicine. Available from: https://doi.org/10.1016/j.amepre.2009.10.040

Smith, M, Zhang, Y, Fainu, HM, Cavadino, A, Zhao, J, Morton, S, Hopkins, D, Carr, H and Clark, T (2024) Socio-environmental factors associated with active school travel in children at ages 6 and 8 years. Transportation Research Interdisciplinary Perspectives. Available from: https://doi.org/10.1016/j.trip.2024.101026

Walker, C. & van Holstein, E. (2025) Training young co-researchers to interview their parents: The transformative potential of intergenerational interviews. Area. Available from: https://doi.org/10.1111/area.12972

How to cite

Shin, H. (2026, July) How well does an algorithm know the way to school? Geography Directions. https://doi.org/10.55203/HOOU2273

Leave a Reply or Comment