Abstract: Astronomy is entering an era in which survey data are growing by orders of magnitude in both volume and complexity. But is our ability to extract physical insight scaling at the same pace? I will explore this question first through the latest progress on quasar identification and redshift determination in Euclid DR1. Slitless spectroscopy poses a substantially harder inference problem than conventional fiber spectroscopy, with spectral overlap, limited wavelength coverage, and line-identification degeneracies producing structured catastrophic failures. We developed the astrophysics-informed machine learning framework AIMS-z to optimize redshift accuracy and reliability simultaneously. For an independent quasar sample, AIMS-z reaches sigma_NMAD~0.0012 with a catastrophic failure rate of only ~1.3%, approaching the precision regime required for cosmological clustering analyses. These results provide a concrete example of how analysis methods can evolve with survey scale: AIMS-z turns a difficult slitless spectroscopy problem into high-precision, low-outlier redshift inference at scale. The same lesson will be important for Roman and CSST.
I will then ask what astrophysical information can be learned directly from large-scale spectroscopic data. Using DESI, we train a large self-supervised spectral representation model and apply its embedding space to similarity retrieval and the identification of local little red dots (LRDs) and other unusual AGN populations. These results point toward a broader foundation-model approach in which reusable representations learned from large unlabeled datasets support multiple scientific tasks. Such methods offer a new route to studying the physical diversity and evolution of AGN populations at scale, turning larger surveys into deeper physical understanding of the Universe.