Reliable AI Sea Ice Forecasts Depend on Sustained Observations

Reliable AI Sea Ice Forecasts Depend on Sustained Observations

A decade ago, dedicated artificial intelligence (AI) sea ice prediction systems barely existed. Today they number in the dozens, covering both poles on timescales ranging from days to seasons. Calculations that once took days to run on a supercomputer can now be executed in minutes on a laptop. Speed and skill are not the same as trustworthiness. But speed and skill are not the same as trustworthiness. The observations needed to build and stress test these systems remain sparse across both poles. Observational Requirements in the Context of AI prediction systems for Sea ice (ORCAS), a new international scientific community, was formed to bridge that gap. ORCAS was established in 2025 as a Scientific Committee on Oceanic Research (SCOR) working group (of which we are all members) and a task team of the World Weather Research Programme’s (WWRP) Polar Coupled Analysis and Prediction for Services (PCAPS) project. We bring together observational scientists, AI developers, and physical modelers from around the world to assess what observations emerging AI-based sea ice prediction systems actually need. Our community is globally distributed and heavily driven by early-career researchers. Observations Still Matter Most Sea ice prediction systems have evolved into practical tools that underpin climate services, operational decisionmaking, and the safety of high-latitude communities. Indigenous communities in the Arctic depend on reliable forecasts for safe travel and hunting. Seagoing vessels need near-real-time predictions to avoid dangerous ice floes. Fisheries, search and rescue teams, and coastal planners also depend on accurate forecasts. Such systems began as research efforts to understand whether sea ice conditions could be predicted from one season to the next [Chen and Yuan, 2004], as well as how far ahead those changes could be predicted [Holland et al., 2013]. In the years since, they have grown into a major forecasting enterprise. Innovative methods, boosted by scientific advances in Arctic sea ice prediction [Bushuk et al., 2024] (Figure 1), now provide near-real-time forecasts at both poles. The growth of international prediction efforts, including the Sub-seasonal to Seasonal (S2S) project [Vitart and Robertson, 2018] and the Sea Ice Prediction Network (SIPN) South [Massonnet et al., 2023], illustrates how quickly the field has moved forward. Fig. 1. This chart displays a yearly count of peer-reviewed publications returned by a Scopus query pairing sea ice prediction terms with AI and machine learning keywords. The field was small and steady through 2018, then accelerated sharply. Two events bracket the steepest rise: the 2021 release of IceNet [Andersson et al., 2021], which showed that deep learning could match or beat dynamical models for seasonal sea ice forecasts, and the 2023 collapse of Antarctic sea ice to record lows. The 2026 hatched bar reflects publications from only January–May. Credit: Created by the authors using Python and Matplotlib from a Scopus query retrieved May 2026 Today the sea ice prediction landscape spans a range of approaches that ORCAS aims to incorporate. Traditional process-based dynamical climate models, for example, simulate interactions between the atmosphere, ocean, and ice. These models remain widely used because they offer extensive hindcast records for testing forecast skill and large ensembles for estimating uncertainty. They can, however, be computationally expensive and prone to systematic biases. Emerging AI-based approaches, including data-driven machine learning models, physics-informed hybrid systems, and large-scale AI weather models, are showing promise as robust and efficient options (Figure 2), often running faster and with a predictive skill that challenges traditional methods. These approaches mostly fall into two camps, offering short-range regional forecasts of a few days to a week for navigation or longer-term pan-Arctic outlooks of ice concentration or extent weeks to months in the future. So far, few focus on Antarctica, and fewer still target conditions at the local scales that matter most for communities and navigation of smaller vessels. Fig. 2. This diagram, drawn from a review of key sea ice models, groups 20 existing systems by what they predict (left), their model class (middle), and how they generate a forecast (right). Most models forecast either sea ice concentration or extent rather than several variables at once, and most produce a forecast by directly predicting a future state rather than stepping forward in time (autoregressive) or processing an input sequence (recurrent). A small number are emulators that reproduce the behavior of physics-based models with long-term stability in mind. The widths of the bands are proportional to the number of systems. Credit: Lorenzo Zampieri Process-based and AI approaches to climate modeling are less independent than they might appear. AI models are often trained on reanalyses derived from process-based simulations, creating a shared lineage. Both rely on observational data for tuning and initialization (process-based models) or model testing and training (AI-based approaches). Near-real-time, reliable, easily accessible, and well-documented observations are the keystone of sea ice prediction, regardless of approach. The Observation Gap Across the polar regions, observational data collection remains sparse, unevenly distributed, and difficult and costly to sustain. As climate change drives both poles into conditions without historical precedent, prediction systems are being asked to perform in scenarios that are not well represented in the observations used to build and test them. AI-ready and process-focused observations offer distinct benefits and challenges for the emerging ORCAS community (Figure 3). AI-ready data are typically standardized, quality controlled, gridded, well-documented, and available at scales suitable for ingestion by machine learning systems (e.g., ERA5 global reanalysis fields [Hersbach et al., 2020]). Process-focused observations, by contrast, are designed to reveal how the sea ice system works, for example, through detailed measurements of ice thickness, drift, and deformation; snow properties; or the exchange of heat between ocean, ice, and atmosphere. These observations may be local, irregular, campaign based, or difficult to harmonize, but they are uniquely valuable for testing whether predictions are physically credible and for providing independent data used for model validation (but not for model training). Fig. 3. At each stage of an AI forecast—training, inference/forecasting, and verification—current systems rely largely on gridded data products, including reanalyses, atmospheric analyses and forecasts, and satellite remote sensing. In situ and campaign measurements (e.g., from buoys, ship-based data, and field campaigns) remain underused at every stage, largely because they are not yet in formats that models can readily ingest, either directly or through reanalyses. Determining how to best strengthen each stage with these sparse, but high-value, observations is a central goal of the Observational Requirements in the Context of AI prediction systems for Sea ice (ORCAS) effort. Credit: Lorenzo Zampieri One general challenge for sea ice prediction systems will be to make observational datasets more accessible, interoperable, and usable for AI training and evaluation—in short, to bring them in line with the FAIR (findable, accessible, interoperable, and reusable) principles. Another challenge is posed by satellite data, which are plentiful for the polar regions. Passive microwave sensors have mapped sea ice concentrations across both poles since 1978, providing the long, consistent records on which most AI sea ice systems are trained and evaluated. Data from synthetic aperture radar, optical imagery, altimetry, and infrared sensors add details about ice types, thicknesses, and deformation. Yet much of these data remain fragmented or difficult to harmonize across missions and time periods. Campaign-based datasets face similar barriers. There is also a scale gap to bridge. Point-scale measurements and campaign observations must connect to the gridded, basin-scale products that AI models typically ingest. Emerging services like IceBox, intended for shipboard decisionmakers at individual sites or along specific routes, are working to close this gap by pushing forecast spatial resolution well below current operational scales, to 100 meters or finer. Building Trust in AI Sea Ice Predictions The rapid adoption of AI prediction systems across the Earth sciences [Eyring et al., 2024] raises a fundamental question: How do we ensure these tools are physically consistent, skillful, and trustworthy? For sea ice prediction, that means testing not only whether a model gives accurate forecasts but also whether its behavior is physically sensible, reliable across different conditions, and useful in the real world. Such evaluations require shared test cases, common datasets, and clear frameworks for transparent benchmarking across AI architectures and against traditional numerical models. Equally important is the ability to forecast out-of-sample events (i.e., those not included in the model’s training set), which have grown increasingly common in recent years. Independent (external) evaluation is central to that effort. Such evaluations require shared test cases, common datasets, and clear frameworks for transparent benchmarking across AI architectures and against traditional numerical models [Ullrich et al., 2025]. Through ORCAS, we aim to develop guidelines for evaluating sea ice prediction systems across forecast timescales. The guidelines will use shared validation scenarios, agreed-upon observational datasets, and metrics that link model performance to real sea ice processes. A key outcome from our recent ORCAS Community Workshop was the conclusion that evaluation must go beyond simple accuracy metrics to incorporate process-focused observations [Eayrs et al., 2026]. These observations must be linked explicitly across scales through targeted metrics (e.g., ice drift, thickness, and transport). AI systems must also be checked against known physical laws and tested for whether a model trained in one region (the Arctic, e.g.) can reliably predict conditions in another (Antarctica). Transparency also matters. AI models are sometimes described as black boxes, producing accurate forecasts without making their reasoning visible or interpretable [Camps-Valls, 2026]. ORCAS will support the use of interpretability tools that reveal which inputs most influence a model’s output [Rudin, 2019]—alongside better uncertainty quantification—so decisionmakers can understand and act on what they are shown. Overall, the evaluation frameworks developed by ORCAS will aim to test whether AI sea ice predictions remain physically credible across different regions, forecast conditions, and a changing climate. Developing these frameworks demands iterative codesign among observational scientists, modelers, and data scientists to ensure physically meaningful benchmarks and relevance for real-world applications. Observations for the AI Age Whereas forecast skill intercomparisons are already well served by initiatives like SIPN South [Massonnet et al., 2023], ORCAS concentrates on a less addressed question: Are AI predictions physically realistic, and can we trust them? To help answer this, we are curating campaign and process study datasets into targeted validation suites. A natural concern is whether prioritizing AI-ready observations risks sidelining process-rich measurements that are harder to standardize but essential for physical understanding. We address this concern through two parallel strategies. First, we are adapting observational priorities so that key variables are collected in forms AI systems can readily use. Second, we are working with developers to incorporate process-focused datasets as evaluation benchmarks that test physical credibility beyond simple accuracy metrics. Our work is organized around three deliverables. The first is a validation exercise comparing AI predictions against historical field campaigns, through which we will develop reusable assessment frameworks. The second is a strategic report on observational priorities to inform planning for the Fifth International Polar Year. The third is an accessible guide for newcomers to the field. Applications Beyond Sea Ice Sea ice sits at one of the most demanding frontiers for AI prediction. Sea ice sits at one of the most demanding frontiers for AI prediction. Data are sparse, conditions are changing faster than models were trained to handle, and the physical processes involved are tightly coupled across scales from individual ice floes to entire ocean basins. But these challenges are not unique to the polar regions. Across the Earth system, AI prediction capabilities are also advancing faster than our collective ability to verify them. The gap between forecast skill and physical trustworthiness, between rapid model development and the slower work of building interpretable, well-grounded evaluation frameworks, is felt across climate prediction, ocean modeling, and atmospheric science. ORCAS offers a response to that challenge, with prediction systems grounded in sea ice but pointing toward something broader. The observations we prioritize, the evaluation frameworks we build, and the standards we develop will matter well beyond the next seasonal forecast. Getting this right in one of the most data-sparse and rapidly changing parts of the planet offers a contribution to the much larger effort of making AI prediction trustworthy across the Earth sciences, one in which we hope the wider community will join us. Acknowledgments The authors thank the members of the ORCAS community—SCOR Working Group 173 and the World Meteorological Organization (WMO) WWRP PCAPS Task Team on Observational Requirements for AI-based Sea Ice Prediction—whose collective discussions shaped the ideas presented here. We are particularly grateful to ORCAS cochair Lorenzo Zampieri (European Centre for Medium-Range Weather Forecasts) for his contributions to the framing of this article and helpful review of the manuscript and for providing figures. We also thank Kristina Dahl (Climate Central) for reviewing the article. We acknowledge the support of SCOR and the WMO WWRP in establishing and sustaining ORCAS. ORCAS remains open to new participants from the observational, modeling, and AI communities. Readers are invited to join the ORCAS community mailing list at https://forms.gle/GGC1HBwypAE6NUNg9, through which we announce annual in-person workshops, held alongside major conferences, and opportunities to contribute models and datasets to the validation exercise. References Andersson, T. R., et al. (2021), Seasonal Arctic sea ice forecasting with probabilistic deep learning, Nat. Commun., 12(1), 5124, https://doi.org/10.1038/s41467-021-25257-4. Bushuk, M., et al. (2024), Predicting September Arctic sea ice: A multimodel seasonal skill comparison, Bull. Am. Meteorol. Soc., 105, E1170–E1203, https://doi.org/10.1175/BAMS-D-23-0163.1. Camps-Valls, G. (2026), AI needs a new philosophy of science, Innovation, 7(5), 101311, https://doi.org/10.1016/j.xinn.2026.101311. Chen, D., and X. Yuan (2004), A Markov model for seasonal forecast of Antarctic sea ice, J. Clim., 17, 3,156–3,168, https://doi.org/10.1175/1520-0442(2004)017%3C3156:AMMFSF%3E2.0.CO;2. Eayrs, C., et al. (2026), The ORCAS Community Workshop: Connecting Observations and AI Sea Ice Prediction Systems—Workshop Report, Zenodo, https://doi.org/10.5281/zenodo.19543215. Eyring, V., et al. (2024), Pushing the frontiers in climate modelling and analysis with machine learning, Nat. Clim. Change, 14(9), 916–928, https://doi.org/10.1038/s41558-024-02095-y. Hersbach, H., et al. (2020), The ERA5 global reanalysis, Q. J. R. Meteorol. Soc., 146, 1,999–2,049, https://doi.org/10.1002/qj.3803. Holland, M. M., et al. (2013), Initial-value predictability of Antarctic sea ice in the Community Climate System Model 3, Geophys. Res. Lett., 40(10), 2,121–2,124, https://doi.org/10.1002/grl.50410. Massonnet, F., et al. (2023), SIPN South: Six years of coordinated seasonal Antarctic sea ice predictions, Front. Mar. Sci., 10, 1148899, https://doi.org/10.3389/fmars.2023.1148899. Rudin, C. (2019), Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead, Nat. Mach. Intell., 1(5), 206–215, https://doi.org/10.1038/s42256-019-0048-x. Ullrich, P. A., et al. (2025), Recommendations for comprehensive and independent evaluation of machine learning–based Earth system models, J. Geophys. Res. Mach. Learn. Comput., 2(1), e2024JH000496, https://doi.org/10.1029/2024JH000496. Vitart, F., and A. W. Robertson (2018), The sub-seasonal to seasonal prediction project (S2S) and the prediction of extreme events, npj Clim. Atmos. Sci., 1, 3, https://doi.org/10.1038/s41612-018-0013-0. Author Information Clare Eayrs ([email protected]), New York University, New York; Zachary Labe, Climate Central Inc., Princeton, N.J.; François Massonnet, Université catholique de Louvain, Louvain-la-Neuve, Belgium; Tian Tian, Danish Meteorological Institute, Copenhagen, Denmark; and Malte Müller, Meteorologisk institutt, Oslo, Norway; also at University of Oslo, Oslo, Norway Citation: Eayrs, C., Z. Labe, F. Massonnet, T. Tian, and M. Müller (2026), Reliable AI sea ice forecasts depend on sustained observations, Eos, 107, https://doi.org/10.1029/2026EO260317. Published on 9 October 2026. Text © 2026. The authors. CC BY 3.0Except where otherwise noted, images are subject to copyright. Any reuse without express permission from the copyright owner is prohibited.

Original Source

Read the full article at Eos →

KhanList aggregates and links to publicly available news content. We do not host full articles from third-party sources. Always verify important information with original sources.