Astronomical Foundation Model Biased by Survey Detection Channel

Ihor Kendiukhov· August 26, 2026 View original

Key takeaways

  • Astronomical foundation models can inherit biases from data pipeline incompleteness.
  • The survey detection channel significantly overrides pixel data, biasing model outputs.
  • This leads to systematic errors in critical measurements like tomographic mean redshifts.
  • Causal auditing is essential to uncover and address such hidden biases in large models.

Who benefits

AstronomyScientific ResearchSpace ExplorationData ScienceAI Ethics

Summary

An audit of AION-1, an astronomical foundation model, reveals that its survey detection channel overrides pixel data, introducing significant biases in reported quantities like redshift. This bias stems from the incompleteness of catalogue products used in training, leading to systematic errors.

A recent audit of AION-1, a large astronomical foundation model trained on over 200 million objects and 39 modalities, has uncovered a critical systematic bias. The model, which learns from both survey pixels and the derived catalogue products, inherits incompleteness from these catalogues. Specifically, the survey detection channel within the model significantly influences its outputs, overriding the raw pixel data. Causal interventions demonstrated that merely editing the survey segmentation map, while keeping image tokens identical, altered every quantity the model reported—including flux, size, ellipticity, and redshift—by factors 110 to 4400 times greater than a placebo. This effect is primarily driven by the presence of an object at the field center, rather than the actual light enclosed by the mask. The model also largely ignores how pipelines partition light in blended objects and performs worse when contradicted by catalogue photometry. A significant consequence is the impact on tomographic mean redshifts. The Legacy Survey pipeline's 3.68% target miss rate, when propagated, shifts these redshifts by a median of 0.71 times the LSST DESC requirement, exceeding it in 12 out of 40 assignments. This bias is not mitigated by magnitude-dependent missing data and grows with model scale. The study suggests that withholding the detection channel can remove this effect without measurable cost, and also points out limitations in the model's tokenizer regarding image resolution and redshift quantization.

Why it matters

Professionals working with large-scale AI models in scientific domains, especially those trained on diverse data sources, must be aware of how data pipeline artifacts and input channel interactions can introduce subtle yet significant biases, impacting the reliability of scientific conclusions.

How to implement this in your domain

  1. 1Conduct thorough causal audits of foundation models to identify hidden biases introduced by specific input channels or data processing steps.
  2. 2Implement rigorous data quality checks for all input modalities, particularly for catalogue products used in training.
  3. 3Explore alternative training strategies that de-emphasize or remove potentially biasing input channels, such as detection maps.
  4. 4Develop methods to quantify and correct for systematic biases in model outputs, especially for critical scientific measurements like redshift.
  5. 5Improve tokenization strategies to ensure high-fidelity representation of both image and spectral data in foundation models.

Original post by Ihor Kendiukhov

"arXiv:2608.23626v1 Announce Type: new Abstract: Foundation models for astronomy are trained on survey pixels together with the catalogue products derived from those pixels. Those catalogues are incomplete at a measurable rate, and a model trained on both inherits that incompleten…"

View on X

Originally posted by Ihor Kendiukhov on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses

More in AI Research

AI ResearchAI Engineering & DevToolsAI Investing

FraudBench Benchmarks Adversarial Robustness in Financial Risk Assessment

This paper introduces FraudBench, a protocol-sensitive benchmark for evaluating the adversarial robustness of machine learning models in financial fraud and credit-risk detection. It demonstrates that robustness conclusions are highly dependent on how domain-specific constraints and attacker capabilities are incorporated into the evaluation protocol.

Xitong Zeng, Zhaoge Bi, Yitian Yang, Huaming Chen, Quan Z. ShengAug 26, 2026
AI ResearchAI Engineering & DevTools

Persistent Cross Entropy Extends Topological Data Analysis

This paper introduces Persistent Cross Entropy (PCE), a novel extension of cross-entropy to persistence diagrams, which are used in topological data analysis. PCE bridges different event spaces of diagrams using an induced probability, enabling new applications like distinguishing diagrams with similar persistent entropy and separating causal directions in dynamical systems.

Sijin Yeom, Jae-Hun JungAug 26, 2026
AI ResearchAI Engineering & DevTools

Bridging Numerical PDE Solvers and Neural Emulators for Faster Simulation

This thesis explores the deep connections between traditional numerical solvers for Partial Differential Equations (PDEs) and neural emulators, arguing that they are more alike than different. It proposes that insights can flow profitably in both directions, leading to faster and more efficient scientific and engineering simulations.

Felix KoehlerAug 26, 2026