ATMOSPHERIC SCIENCE · DATA ASSIMILATION
Two experiments. One question: when and where does Data Assimilation matter most?
01 · INTRODUCTION
Heatwaves are prolonged periods of abnormally high temperatures that can cause severe human and economic damage. As climate change accelerates, heatwaves are becoming more frequent and intense, making accurate forecasting increasingly important.
This study compares two dimensions: time (2021 vs 2024) and space (Seoul/Gyeonggi vs Gangwon) to investigate how DA performs under different heatwave intensities and observation network densities.
02 · DA THEORY
DA is a mathematical framework that optimally combines a model background with real-world observations — giving more trust to whichever source is more reliable — to produce an improved estimate of the true atmospheric state called the analysis field.
Dense observation networks allow DA to constrain the analysis field more tightly, reducing background errors more effectively. In sparse regions, the system relies more heavily on the model background — leaving initial condition errors less corrected.
This study directly tests this by comparing Seoul/Gyeonggi (SE) — observation-dense — against Gangwon (GW) — observation-sparse — during the 2024 heatwave.
03 · METHOD & TEAM
The team of four divided responsibilities across data pipeline, core algorithm, temporal analysis, and spatial analysis — each building on a shared EnKF framework.
04 · RESULTS
Results from two experiments: time comparison (2021 vs 2024) and spatial comparison (SE vs GW during 2024 heatwave).
Comparing DA performance between normal year (2021) and heatwave year (2024) across SE and GW regions.
05 · CONCLUSION
When atmospheric conditions are most extreme and observation networks are most sparse, observational data becomes most valuable — but its impact is fundamentally limited by how many observations are available.
06 · LIMITATIONS
07 · EXTENDED INTERPRETATION
Both experiments point to the same underlying cause: ERA5 has fundamentally different bias characteristics in urban vs. mountainous regions.
ERA5's ~31km grid averages Seoul's dense urban core with surrounding suburban areas, systematically overestimating temperature. During the 2024 widespread heatwave, the urban–suburban gradient narrowed, reducing ERA5's overestimation bias.
Gangwon's Taebaek range produces sharp elevation gradients and föhn-driven temperature surges that ERA5 cannot resolve. During the 2024 heatwave, these localized extremes intensified, widening the background error.
Key Insight: The opposing increment directions and asymmetric RMSE trends both reflect the same structural truth — ERA5 overestimates in urban regions and underestimates in mountainous terrain. DA corrects these biases, but magnitude and direction differ by region.
08 · REAL-WORLD IMPLICATIONS
Our three key limitations are exactly what KMA's operational Korean Integrated Model (KIM) has been built to solve.
KIM runs a 6-hour cycled 4DEnVar system — each analysis feeds into the next cycle, creating true temporal continuity.
KIM blends ensemble and static background error covariance — when ensemble spread collapses, the static component compensates.
KIM applies tailored observation operators for each data type with spatial interpolation and bias correction. The 8km upgrade further reduced grid-observation mismatch.