Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Data Catalog”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Short GRB Host Galaxies. I. Photometric and Spectroscopic Catalogs, Host Associations, and Galactocentric Offsets

We present a comprehensive optical and near-infrared census of the fields of 90 short gamma-ray bursts (GRBs) discovered in 2005–2021, constituting all short GRBs for which host galaxy associations are feasible (≈60% of the total Swift short GRB population). We contribute 274 new multi-band imaging observations across 58 distinct GRBs and 26 spectra of their host galaxies. Supplemented by literature and archival survey data, the catalog contains 542 photometric and 42 spectroscopic data sets. The photometric catalog reaches 3σ depths of ≳24–27 mag and ≳23–26 mag for the optical and near-infrared bands, respectively. We identify host galaxies for 84 bursts, in which the most robust associations make up 56% (50/90) of events, while only a small fraction, 6.7%, have inconclusive host associations. Based on new spectroscopy, we determine 18 host spectroscopic redshifts with a range of z ≈ 0.15–1.5 and find that ≈23%–41% of Swift short GRBs originate from z > 1. We also present the galactocentric offset catalog for 84 short GRBs. Taking into account the large range of individual measurement uncertainties, we find a median of projected offset of ≈7.7 kpc, for which the bursts with the most robust associations have a smaller median of ≈4.8 kpc. Our catalog captures more high-redshift and low-luminosity hosts, and more highly offset bursts than previously found, thereby diversifying the population of known short GRB hosts and properties. In terms of locations and host luminosities, the populations of short GRBs with and without detectable extended emission are statistically indistinguishable. This suggests that they arise from the same progenitors, or from multiple progenitors, which form and evolve in similar environments. All of the data products are available on the Broadband Repository for Investigating Gamma-Ray Burst Host Traits website.

79 ASTRONOMY AND ASTROPHYSICS↗

Utah FORGE: Slide-Hold-Slide Experiments on Gneiss at Increased Temperature

Included are data from triaxial, single-inclined-fracture friction experiments. The experiments were performed with slide-hold-slide protocol on Utah FORGE gneiss at increased temperature. With a ~10 MPa normal stress, temperatures vary between experiments from room temperature up to 163 Celsius. Hold times vary during experiment from ~10^1 to ~10^5 seconds. Measured are the frictional response upon reactivation after a hold period, active acoustic data (P-wave velocity and amplitude) and passive acoustic data (acoustic emission occurrence and amplitude). There are two types of datafiles: (1) Datafiles containing the friction data, including the temperature and the active acoustic data measured during the experiment (AEXX_Gneiss_Vp_mixref4). The underscore _Vp means that it includes the Vp or P-wave velocity data, with _mixref meaning that we use a mixed reference point for calculating the P-wave velocity. And (2) the datafiles containing the passive acoustics data, a catalog of the acoustic emissions (AE's) measured during the experiment (AEcatalog_AEXX_runX), where AEXX matches the experiment number and runX denotes which part of the experiment the data was collected, matching the times where active acoustic data was collected. AE catalogs are split in two parts when the file size exceeds 1 GB to aid download/opening times.

15 GEOTHERMAL ENERGY↗

A Catalog of Quasar Properties from Sloan Digital Sky Survey Data Release 16

We present a catalog of continuum and emission-line properties for 750,414 broad-line quasars included in the Sloan Digital Sky Survey Data Release 16 quasar catalog (DR16Q), measured from optical spectroscopy. These quasars cover broad ranges in redshift (0.1 ≲ z ≲ 6) and luminosity (44 ≲ log(L bol /erg s -1 ) ≲ 48), and probe lower luminosities than an earlier compilation of SDSS DR7 quasars. Derived physical quantities such as single-epoch virial black hole masses and bolometric luminosities are also included in this catalog. We present improved systemic redshifts and realistic redshift uncertainties for DR16Q quasars using the measured line peaks and correcting for velocity shifts of various lines with respect to the systemic velocity. About 1%, 1.4%, and 11% of the original DR16Q redshifts deviate from the systemic redshifts by |ΔV| > 1500 km s -1 , |ΔV| $\in$ [1000, 1500] km s -1 , and |ΔV| $\in$ [500, 1000] km s -1 , respectively; about 1900 DR16Q redshifts were catastrophically wrong (|ΔV| > 10,000 km s -1 ). We demonstrate the utility of this data product in quantifying the spectral diversity and correlations among physical properties of quasars with large statistical samples.

79 ASTRONOMY AND ASTROPHYSICS↗

Requirements for Cataloging Hanford Geophysical Datasets

Environmental management activities at the Hanford Site produce extensive data about site conditions, contaminants, cleanup, and more. Managing and archiving that data requires a high degree of collaboration among site contractors and a high level of awareness by project managers and staff. Part of that effort is developing a Hanford Environmental Information and Data Index (HEIDI) to organize the data and maximize its value by making it findable and available for reuse. The objective is to catalog the disparate data sets collected to address the evolving needs of planning, executing, and documenting cleanup over several decades up to the present day, including links to active data sources when available. A properly implemented data catalog makes finding environmental datasets related to an area or theme a routine, reliable process, without requiring the searcher to have special knowledge that a data set exists and where it may be stored. In this project, a working group, including the U.S. Department of Energy, the Hanford Site contractors, and Pacific Northwest National Laboratory staff, identified needs and requirements for handling complex site data. Geophysical data was chosen as a test case because it can be large and complex and often involves multiple processing steps to extract the information incorporated into deliverables. The ability to document those steps was one of the requirements identified for the catalog. In addition to developing requirements, other activities included selecting a metadata schema and initial testing with the objective of determining whether the workflow and capabilities of selected data catalog software platforms were sufficient to implement and impose the identified requirements. This initial testing involved running the default catalog instance using the software platform of interest and altering the configuration to achieve each requirement, if possible. Where configuration alone was insufficient, the possibility of modifying the software by changing the code was examined, but not implemented. A follow-on task is planned to reprogram the code as necessary to implement requirements in a prototype catalog.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

Evolution of DUNE’s Production System

The DUNE experiment will start running in 2029 and record 30 PB/year of raw waveforms from Liquid Argon TPCs and photon detectors. The size of individual readouts can range from 100 MB to a typical 8 GB full readout of the detector, and even 100 TB for extended readouts from supernova candidates. These data then need to be cataloged, stored and distributed for processing worldwide. This massive amount of data and a heterogeneous computing environment necessitates a powerful and robust distributed computing infrastructure. In the process of building up that infrastructure, DUNE’s production system has recently undergone an overhaul, in which it has integrated 1) a new workflow management system (justIN) 2) a new data catalog (MetaCat) and 3) a state-of-the-art data management system (Rucio). Simulations of DUNE’s Far Detector and its prototypes ProtoDUNE Horizontal Drift (ProtoDUNE-HD) and ProtoDUNE Vertical Drift (ProtoDUNE-VD), as well as data from ProtoDUNE-HD serve as the first tests of this infrastructure.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Deliverable D11 – Data Sharing, Storage, Security Protocols, and a Specification of a Potential Data Sharing Portal

Pacific Northwest National Laboratory (PNNL) and Technical University of Denmark (DTU) completed this deliverable as part of Work Package 2: Data Information Catalog for Distributed Wind Research for the International Energy Agency (IEA) Wind Technology Collaboration Programme Task 41: Enabling Wind to Contribute to a Distributed Energy Future. As part of the work plan, Deliverable D11 requires the development of data sharing, storage, and security protocols for metadata to be stored on the platform, if needed. The specification of a potential data sharing portal that expands on the catalog is also required.

17 WIND ENERGY↗

Ultrahigh-resolution mass spectrometry data associated with the manuscript “A functional microbiome catalog crowdsourced from North American rivers"

This data package is associated with the publication “A functional microbiome catalog crowdsourced from North American rivers” submitted to Nature (Borton et al., 2024); (https://www.biorxiv.org/content/10.1101/2023.07.22.550117v1). Predicting elemental cycles and maintaining water quality under increasing anthropogenic influence requires understanding the spatial drivers of river microbiomes. However, the unifying microbial determinants governing river biogeochemistry are hindered by a lack of genome-resolved functional insights and sampling across multiple rivers. Here we employed a community science effort to accelerate the sampling of river microbiomes to create the Genome Resolved Open Watersheds database (GROWdb). GROWdb is a publicly available resource that paves the way for watershed predictive modeling and microbiome-based management practices. This resource profiled the identity, distribution, function, and expression of thousands of microbial genomes across rivers covering 90% of United States watersheds. We identified the most cosmopolitan microbiome members, while also revealing local drivers of strain endemism across ecological dimensions. We provide the first evidence that microbial functional trait expression followed the tenets of the River Continuum Concept, suggesting the structure and function of river microbiomes is predictable. The Fourier-transform ion cyclotron resonance mass spectrometry (FTICR-MS) data were one of many different data types used in establishing the ecological dimensions along which different microbes were detected .This data package only contains the processed FTICR-MS data associated with this manuscript; all other data is accessible via Zenodo (https://zenodo.org/records/8173287), GitHub (https://github.com/jmikayla1991/Genome-Resolved-Open-Watersheds-database-GROWdb), KBase (https://doi.org/10.25982/109073.30/1895615), and NCBI via Bioproject PRJNA946291.This dataset consists of (1) a file-level metadata (flmd) file; (2) a data dictionary (dd) file; (3) a readme; (4) three Fourier-transform ion cyclotron resonance mass spectrometry (FTICR-MS) processed data files (a ‘data’ file containing peak-by-sample observations, a ‘mol’ file containing peak metadata, and a transformation profile containing transformation-by-sample observations). All files are .csv or .pdf.

54 ENVIRONMENTAL SCIENCES↗

Millimeter-wave observations of Euclid Deep Field South using the South Pole Telescope: A data release of temperature maps and catalogs

Context. The South Pole Telescope third-generation camera (SPT-3G) has observed over 10,000 square degrees of sky at 95, 150, and 220 GHz (3.3, 2.0, 1.4 mm, respectively) and will significantly overlap the ongoing 14,000 square-degree Euclid Wide Survey. The Euclid collaboration recently released Euclid Deep Field South (EDF-S) observations of 23 square degrees at wide field depths in the first quick data release (Q1). Aims. With the goal of releasing complementary millimeter-wave data and encouraging legacy science, we performed dedicated observations of a 57-square-degree field overlapping the EDF-S. Methods. The observing time totaled 20 days, and we reached noise depths of 4.3, 3.8, and 13.2 $μ$K-arcmin at 95, 150, and 220 GHz, respectively. Results. In this work we present the temperature maps and two catalogs constructed from these data. The emissive source catalog contains 601 objects (334 inside EDF-S) with 54% synchrotron-dominated sources and 46% thermal dust emission-dominated sources. The 5$σ$ detection thresholds are 1.7, 2.0, and 6.5 mJy in the three bands. The cluster catalog contains 217 cluster candidates (121 inside EDF-S) with median mass $M_{500c}=2.12 \times 10^{14} M_{\odot}/h_{70}$ and median redshift $z$ = 0.70, corresponding to an order-of-magnitude improvement in cluster density over previous tSZ-selected catalogs in this region (3.81 clusters per square degree). Conclusions. The overlap between SPT and Euclid data will enable a range of multiwavelength studies of the aforementioned source populations. This work serves as the first step toward joint projects between SPT and Euclid and provides a rich dataset containing information on galaxies, clusters, and their environments.

Archipley, M. [Chicago U., Astron. Astrophys. Ctr.↗

Opportunities in AI/ML for the Rubin LSST Dark Energy Science Collaboration

The Vera C. Rubin Observatory's Legacy Survey of Space and Time (LSST) will produce unprecedented volumes of heterogeneous astronomical data (images, catalogs, and alerts) that challenge traditional analysis pipelines. The LSST Dark Energy Science Collaboration (DESC) aims to derive robust constraints on dark energy and dark matter from these data, requiring methods that are statistically powerful, scalable, and operationally reliable. Artificial intelligence and machine learning (AI/ML) are already embedded across DESC science workflows, from photometric redshifts and transient classification to weak lensing inference and cosmological simulations. Yet their utility for precision cosmology hinges on trustworthy uncertainty quantification, robustness to covariate shift and model misspecification, and reproducible integration within scientific pipelines. This white paper surveys the current landscape of AI/ML across DESC's primary cosmological probes and cross-cutting analyses, revealing that the same core methodologies and fundamental challenges recur across disparate science cases. Since progress on these cross-cutting challenges would benefit multiple probes simultaneously, we identify key methodological research priorities, including Bayesian inference at scale, physics-informed methods, validation frameworks, and active learning for discovery. With an eye on emerging techniques, we also explore the potential of the latest foundation model methodologies and LLM-driven agentic AI systems to reshape DESC workflows, provided their deployment is coupled with rigorous evaluation and governance. Finally, we discuss critical software, computing, data infrastructure, and human capital requirements for the successful deployment of these new methodologies, and consider associated risks and opportunities for broader coordination with external actors.

Aubourg, Eric [APC, Paris] (ORCID:000000025592023X↗

Phase curves of small bodies from the SLOAN Moving Objects Catalog

Extensive photometric surveys continue to produce enormous stores of data on small bodies. These data are typically sparsely obtained at arbitrary (or unknown) rotational phases. Therefore, new methods for processing such data need to be developed to make the most of these vast catalogs. We aim to produce a method of recreating the phase curves of small bodies by considering the uncertainties introduced by the nominal errors in the magnitudes and the effect introduced by rotational variations. Here, we use the SLOAN Moving Objects Catalog data as a benchmark to construct phase curves of all small bodies in u', g', r', i', and z' filters. From the phase curves, we obtain the absolute magnitudes and we use them to set up the absolute colors, which are the colors of the asteroids that are not affected by changes in the phase angle. We selected objects with ≥3 observations taken in at least one filter and spanning over a minimum of 5 degrees in the phase angle. We developed a method that combines Monte Carlo simulations and Bayesian inference to estimate the absolute magnitudes using the HG 12 * photometric system. We obtained almost 15 000 phase curves, with about 12 000 of these including all five filters. The absolute magnitudes and absolute colors are compatible with previously published data that support our method. The method we developed is fully automatic and well suited for a run based on large amounts of data. Moreover, it includes the nominal uncertainties in the magnitudes and the whole distribution of possible rotational states of the objects producing what are possibly less precise values, that is, larger uncertainties, but more accurate, namely, closer to the actual value. To our knowledge, this work is the first to include the effect of rotational variations in such a manner.

79 ASTRONOMY AND ASTROPHYSICS↗

Constructing a High‐Resolution Aftershock Catalog for the 2017 Mw 8.2 Tehuantepec Earthquake Sequence Using a Machine Learning–Based Workflow

The 8 September 2017 Mw 8.2 Tehuantepec earthquake was the largest instrumentally recorded normal‐faulting earthquake in Mexico. The mainshock occurred offshore within the Tehuantepec seismic gap, generating >30,000 aftershocks in the following year. We applied an open‐source, machine learning (ML)–assisted workflow to construct a high‐resolution aftershock catalog using data from temporary and permanent seismic networks in southern Mexico. The workflow integrates PhaseNet for phase detection; GaMMA for phase association; and VELEST, HypoInverse, and HypoDD for velocity modeling and relocation. We processed seven months of continuous waveform data from 29 broadband stations, including a temporary rapid‐response deployment that improved station coverage of the offshore rupture zone. To evaluate performance, we compared our results against analyst‐reviewed picks and event locations from the Servicio Sismológico Nacional catalog. The resulting catalog contains 11,374 relocated earthquakes and represents the most comprehensive published dataset for this sequence, incorporating the first full use of the temporary network. Relocated hypocenters show improved depth control and align well with the Slab2.0 subduction geometry, revealing clearer separation between offshore slab events and onshore crustal seismicity. This study demonstrates that combining ML‐based detection with established methods provides a scalable and reproducible approach for constructing high‐quality earthquake catalogs in tectonically complex environments and offers practical guidance for adapting similar workflows to other earthquake sequences.

Garcia, Marc [The University of Texas at El Paso, ↗

DESIVAST: Catalogs of Low-redshift Voids Using Data from the DESI Data Release 1 Bright Galaxy Survey

We present three separate void catalogs created using a volume-limited sample of the DESI Data Release 1 Bright Galaxy Survey. We use the algorithms VoidFinder and V 2 to construct void catalogs out to a redshift of z = 0.24. Excluding voids affected by the boundaries of the survey, we obtain 1489 voids with VoidFinder, 389 with V 2 using REVOLVER pruning, and 297 with V 2 using VIDE pruning. Comparing our catalogs with overlapping Sloan Digital Sky Survey void catalogs, we find generally consistent void properties but significant differences in the void volume overlap, which we attribute to differences in the galaxy selection and survey masks. These catalogs are suitable for studying the variation in galaxy properties with cosmic environment and for cosmological studies.

79 ASTRONOMY AND ASTROPHYSICS↗

Clustering with general photo- z uncertainties: application to Baryon Acoustic Oscillations

ABSTRACT Photometric data can be analysed using the 3D correlation function ξp to extract cosmological information via e.g. measurement of the Baryon Acoustic Oscillations (BAO). Previous studies modeled ξp assuming a Gaussian photo-z approximation. In this work we improve the modeling by incorporating realistic photo-z distribution. We show that the position of the BAO scale in ξp is determined by the photo-z distribution and the Jacobian of the transformation. The latter diverges at the transverse scale of the separation s⊥, and it explains why ξp traces the underlying correlation function at s⊥, rather than s, when the photo-z uncertainty σz/(1+ z) ≳ 0.02. We also obtain the Gaussian covariance for ξp. Due to photo-z mixing, the covariance of ξp shows strong off-diagonal elements. The high correlation of the data causes some issues to the data fitting. None the less, we find that either it can be solved by suppressing the largest eigenvalues of the covariance or it is not directly related to the BAO. We test our BAO fitting pipeline using a set of mock catalogs. The data set is dedicated for Dark Energy Survey Year 3 (DES Y3) BAO analyses and includes realistic photo-z distributions. The theory template is in good agreement with mock measurement. Based on the DES Y3 mocks, ξp statistic is forecast to constrain the BAO shift parameter α to be 1.001 ± 0.023, which is well consistent with the corresponding constraint derived from the angular correlation function measurements. Thus, ξp offers a competitive alternative for the photometric data analyses.

79 ASTRONOMY AND ASTROPHYSICS↗

A Data Processing Pipeline To Extract A Knowledge Graph From Heterogeneous Data For Socio-technical Analysis Of Critical Infrastructure Influence

The code is written in Python and consists of the following pipeline that is implemented in Apache Airflow. This pipeline intends to understand the companies that are directly or indirectly involved with a type of critical infrastructure system at some point in that system's lifecycle. The pipeline takes a configuration file that specifies a list of initial companies to consider, a geographic region of interest, and a set of SEC form types as well as other data sources (e.g. CrunchBase) from which to extract entities and relations. There are four main components to this pipeline as currently implemented: Entity Extraction, Network Construction, Analysis, and Visualization. First, Entity Extraction, is implemented as the `topear-extract_organizations` Apache Airflow workflow. Given an initial query that specifies a geographic region of interest and a time interval, the software will extract CI facilities of interest and organizations that have a direct influence relationship to those facilities (e.g. ownership). During the course of the LDRD, we focused on Electric Vehicle charging stations and this information is available via the Department of Energy (DOE) database on fueling stations maintained by NREL. Within the context of the DOE CESER project, we have focused on Battery Energy Storage Systems (BESS). Second, the Network Extraction component will iteratively construct a social network graph given the set of organizations and people extracted in the previous step. Organizations (and eventually People if desired) are then fed as a query to the `topgear-construct_social_network` Apache Airflow workflow which given a set of initial companies and data sets (e.g. SEC EDGAR form types, OpenCorporates, Crunchbase). This Airflow workflow will iteratively query such data sources to discover relationships with new organizations and people. For example, this module can iteratively query SEC EDGAR for metadata that documents the number of each type of form for the given set of companies and their location. This forms metadata represents a catalog of data sources from SEC EDGAR for the extracted social network knowledge graph. The pipeline then downloads these forms from the website and saves them in a build directory for further processing. These documents are then parsed for entities and relations. Again, we note that in additional to SEC data sources, this step can also pull in information on organizations via API services such as CrunchBase and OpenCorporates or bulk data sources. At the end of this step, the resultant social network, the Critical Infrastructure network, and the edges that encode relationships between organizations and CI facilities, form the Adversarial Socio-Technical Network (ASTN) that informs the analysis. Third, the Analysis component processes these generated ASTN. Previously, that has included the ability to compare prevalence of different vendors for a given infrastructure component type across different regions as well as identify common public and private investors across those vendors. This was demonstrated for EV Charging Stations across several different metropolitan areas within an IEEE PES GridEdge publication. More recently, we have looked at ways to identify infrastructure owners and operators of BESS with the most nameplate capacity across different states as well as other indictors of risk resulting from changes in ownership over time. Finally, the Visualization component consists of an HTML/CSS/JS framework by which users can interact geospatial, operational, and organizational relationships across a given portfolio of Critical Infrastructure facilities. The objective is to provide a library of UI/UX modules that can be repurposed for stakeholder-specific dashboards. All of the modules are related via a common event model that enables UI actions in one view to percolate across the other views.

Weaver, Gabriel [Idaho National Laboratory (INL), ↗

BASS. XXIV. The BASS DR2 Spectroscopic Line Measurements and AGN Demographics

We present the second catalog and data release of optical spectral line measurements and active galactic nucleus (AGN) demographics of the BAT AGN Spectroscopic Survey, which focuses on the Swift-BAT hard X-ray detected AGNs. We use spectra from dedicated campaigns and publicly available archives to investigate spectral properties of most of the AGNs listed in the 70 month Swift-BAT all-sky catalog; specifically, 743 of the 746 unbeamed and unlensed AGNs (99.6%). We find a good correspondence between the optical emission line widths and the hydrogen column density distributions using the X-ray spectra, with a clear dichotomy of AGN types for N H = 10 22 cm –2 . Based on optical emission-line diagnostics, we show that 48%–75% of BAT AGNs are classified as Seyfert, depending on the choice of emission lines used in the diagnostics. The fraction of objects with upper limits on line emission varies from 6% to 20%. Roughly 4% of the BAT AGNs have lines too weak to be placed on the most commonly used diagnostic diagram, [O iii ]λ5007/Hβ versus [N ii ]λ6584/Hα, despite the high signal-to-noise ratio of their spectra. This value increases to 35% in the [O iii ]λ5007/[O ii ]λ3727 diagram, owing to difficulties in line detection. Compared to optically selected narrow-line AGNs in the Sloan Digital Sky Survey, the BAT narrow-line AGNs have a higher rate of reddening/extinction, with Hα/Hβ > 5 (~36%), indicating that hard X-ray selection more effectively detects obscured AGNs from the underlying AGN population. Finally, we present a subpopulation of AGNs that feature complex broad lines (34%, 250/743) or double-peaked narrow emission lines (2%, 17/743).

79 ASTRONOMY AND ASTROPHYSICS↗

A Data Processing Pipeline To Extract A Knowledge Graph From Sec Documents For Socio-technical Analysis Of Critical Infrastructure Influence

The code is written in Python and consists of the following pipeline that is implemented in Apache Airflow. This pipeline intends to understand the companies that are directly or indirectly involved with a type of critical infrastructure system at some point in that system's lifecycle. The pipeline takes a configuration file that specifies a list of initial companies to consider, a geographic region of interest (disk) expressed as a latitude/longitude point and distance, and a set of SEC form types from which to extract entities and relations. There are three main components to this pipeline as currently implemented: Social Network Extraction, Critical Infrastructure Network Extraction, and Inference and Fusion. First, Social Network Extraction, implemented as the `organizations_sec` component of the workflow graph queries the SEC EDGAR webservice using the list of initial companies from the configuration file. Given this, it extracts metadata that documents the number of each type of form for the given set of companies and their location. This forms metadata represents a catalog of data sources for the extracted social network knowledge graph. The pipeline then downloads these forms from the website and saves them in a build directory for further processing. These documents are then parsed for entities and relations. Second, the Critical Network Extraction component extracts entities and relations for a critical infrastructure sector. Currently, we focus on Electric Vehicle charging stations and this information is available via the Department of Energy (DOE) database on fueling stations maintained by NREL. Third, the Inference and Fusion component relates the social network graph to the critical infrastructure graph in order to understand the impact of a company within a geographic region. Relations include ownership of the EV Charging Station asset as well as maintenance/ownership of the EV payment networks. The fused network can be represented in many ways and currently we emit a knowledge graph.

Weaver, GabrielA.↗

A novel cosmic filament catalogue from SDSS data

Here, in this work, we present a new catalogue of cosmic filaments obtained from the latest Sloan Digital Sky Survey (SDSS) public data. In order to detect filaments, we implement a version of the Subspace-Constrained Mean-Shift algorithm that is boosted by machine learning techniques. This allows us to detect cosmic filaments as one-dimensional maxima in the galaxy density distribution. Our filament catalogue uses the cosmological sample of SDSS, including Data Release 16, and therefore inherits its sky footprint (aside from small border effects) and redshift coverage. In particular, this means that, taking advantage of the quasar sample, our filament reconstruction covers redshifts up to z = 2.2, making it one of the deepest filament reconstructions to our knowledge. We follow a tomographic approach and slice the galaxy data in 269 shells at different redshift. The reconstruction algorithm is applied to 2D spherical maps. The catalogue provides the position and uncertainty of each detection for each redshift slice. The quality of our detections, which we assess with several metrics, show improvement with respect to previous public catalogues obtained with similar methods. We also detect a highly significant correlation between our filament catalogue and galaxy cluster catalogues built from microwave observations of the Planck Satellite and the Atacama Cosmology Telescope.

79 ASTRONOMY AND ASTROPHYSICS↗

SpecDis: Value Added Distance Catalog for 4 Million Stars from DESI Year-1 Data

We present the SpecDis value-added stellar distance catalog accompanying DESI Data Release 1. SpecDis trains a feed-forward neural network (NN) with Gaia parallaxes and gets the distance estimates. To build up an unbiased training sample, we do not apply selections on parallax error or signal-to-noise (S/N) of the stellar spectra, and instead, we incorporate parallax error into the loss function. Moreover, we employ principal component analysis to reduce the noise and dimensionality of stellar spectra. Validated by independent external samples of member stars with precise distances from globular clusters, dwarf galaxies, stellar streams, combined with blue horizontal branch stars, we demonstrate that our distance measurements show no significant bias up to 100 kpc, and are much more precise than Gaia parallax beyond 7 kpc. The median distance uncertainties are 23%, 19%, 11%, and 7% for S/N < 20, 20 ≤ S/N < 60, 60 ≤ S/N < 100, and S/N ≥ 100. Selecting stars with ${\mathrm{log}}\,g\lt 3.8$ and distance uncertainties smaller than 25%, we have more than 74,000 giant candidates within 50 kpc of the Galactic center and 1500 candidates beyond this distance. Additionally, we develop a Gaussian mixture model to identify unresolvable equal-mass binaries by modeling the discrepancy between the NN-predicted and the geometric absolute magnitudes from Gaia parallaxes and identify 120,000 equal-mass binary candidates. Our final catalog provides distances and distance uncertainties for >4 million stars, offering a valuable resource for Galactic astronomy.

astronomy data analysis↗