Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “open data format”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16

HydraGNN_Predictive_GFM_2026 - Ensemble of predictive graph foundation models for atomistic materials modeling

This release contains data and parameters of HydraGNN-based graph foundation models trained as a result of the work published in the pre-print "Exascale Multi-Task Graph Foundation Models for Imbalanced, Multi-Fidelity Atomistic Data" by M. Lupo Pasini et al. (https://arxiv.org/abs/2604.15380). We jointly train on 16 open first-principles datasets (544+ million structures covering 85+ elements) using a multi-task architecture with per-dataset heads and a scalable ADIOS2/DDStore data pipeline. On Frontier, we execute six large-scale DeepHyper hyperparameter optimization campaigns in FP64 and promote the top-performing message-passing models to sustained 2,048-node training, yielding a PaiNN-based lead model. The version of HydraGNN used to generate the outputs provided in this release is HydraGNN v5.0 (https://github.com/ORNL/HydraGNN/releases/tag/v5.0) The list of datasets used for the training of the graph foundation model is the following: 1) Alexandria [1] 2) ANI1x [2] 3) MPTrj [3] 4) Open Catalyst 2020 (OC20) [4] 5) Open Catalyst 2022 (OC22) [5] 6) Open Catalyst 2025 (OC25) [6] 7) Open Direct ir Capture 2023 (ODAC23) [7] 8) Open Materials 2024 (OMat24) [8] 9) Open Molecules 2025 (OMol25) [9] 10) OMol25-neutral (subset of OMol25 that contains only molecules with zero total charge) 11) OMol25-non-neutral (subset of OMol25 that contains only molecules with non-zero total charge) 12) Open Polymers 2026 (OPoly2026) [10] 13) Nabla2DFT [11] 14) QCML [12] 15) QM7X [reference 13] 16) transition1x [14] Dataset references: [1] J. Schmidt et al., “A dataset of 175k stable and metastable materials calculated with the PBEsol and SCAN functionals,” Scientific Data, vol. 9, p. 64, 2022. [2] J. S. Smith et al., “The ANI-1ccx and ANI-1x data sets, coupled-cluster and density functional theory properties for molecules,” Scientific Data, vol. 7, p. 134, 2020. [Online]. Available: https: //www.nature.com/articles/s41597-020-0473-z [3] A. Jain et al., “Commentary: The Materials Project: A materials genome approach to accelerating materials innovation,” APL Materials, vol. 1, no. 1, p. 011002, 07 2013. [Online]. Available: https://doi.org/10.1063/1.4812323 [4] L. Chanussot et al., “Open catalyst 2020 (oc20) dataset and community challenges,” ACS Catalysis, vol. 11, no. 10, pp. 6059–6072, 2021. [Online]. Available: https://doi.org/10.1021/acscatal.0c04525 [5] K. Tran et al., “Open catalyst 2022 (oc22) dataset and challenges for oxidation electrocatalysts,” ACS Catalysis, vol. 13, no. 5, pp. 3066–3084, 2023. [Online]. Available: https://doi.org/10.1021/acscatal.2c05426 [6] S. J. Sahoo et al., “The open catalyst 2025 (oc25) dataset and models for solid-liquid interfaces,” arXiv preprint arXiv:2509.17862, 2025. [Online]. Available: https://arxiv.org/abs/2509.17862 [7] A. Sriram et al., “The open DAC 2023 dataset and challenges for sorbent discovery in direct air capture,” ACS Central Science, vol. 10, no. 5, pp. 923–941, 2024. [8] L. Barroso-Luque et al., “Open materials 2024 (omat24) inorganic materials dataset and models,” 2024. [Online]. Available: https://arxiv.org/abs/2410.12771 [9] D. S. Levine et al., “The open molecules 2025 (OMol25) dataset, evaluations, and models,” 2025. [Online]. Available: https://arxiv.org/abs/2505.08762 [10] D. S. Levine et al., The open polymers 2026 (OPoly26) dataset and evaluations,” arXiv preprint arXiv:2512.23117, 2025. [Online]. Available: https://arxiv.org/abs/2512.23117 [11] K. Khrabrov et al., “Nabla2dft: A universal quantum chemistry dataset of drug-like molecules and a benchmark for neural network potentials,” in NeurIPS 2024 Datasets and Benchmarks Track, 2024. [Online]. Available: https://openreview.net/forum?id=ElUrNM9U8c [12] S. Ganscha et al., “The QCML dataset, quantum chemistry reference data from 33.5M DFT and 14.7B semi-empirical calculations,” Scientific Data, vol. 12, p. 406, 2025. [13] J. Hoja et al., “QM7-X, a comprehensive dataset of quantum-mechanical properties spanning the chemical space of small organic molecules,” Scientific Data, vol. 8, p. 43, 2021. [Online]. Available: https://www.nature.com/articles/s41597-021-00812-2 [14] M. Schreiner et al., “Transition1x - a dataset for building generalizable reactive machine learning potentials,” Scientific Data, vol. 9, p. 779, 2022. The folder "datasets_ADIOS2_format" contains the set of pre-processed datasets in Adaptable I/O System (ADIOS) format (https://www.exascaleproject.org/research-project/adios/) that have been used for the development and training of GFMs in this work. The "datasets_ADIOS2_format" directory contains 2 sub-directories, one for the version "v1" of the datasets and one for the version "v2" of the datasets. The version "v1" of the datasets provides values of the total energy as they are extracted from the original data as it was released by the respective institutions. The version "v2" of the datasets provides values of the energy that have been realigned. The realignment was performed by training a linear regression model that predicts the total energy as a function of the chemical composition of the atomistic structure, and then subtract such prediction from the original value of the total energy. Both folders "v1" and "v2" contain 16 sub-directories, each corresponding to an ADIOS2-formatted dataset The folder "DeepHyper-results" contains the configurational files and model's parameters for all the 186 HPO trials that were successfully completed by the scalable hyperparameter optimization (HPO) runs on Frontier. The content of the folder "DeepHyper-results" I structured as follows: 1) task-list.txt: list of mpnn name, jobid, and deephyper task id 2) gfm_${MPNN}_${JOBID}_0.${TASKID}: run directory with checkpoint files 3) gfm_${MPNN}: deephyper summary directory (*.csv) for each specific MPNN type 4) deephyper-experiment-${JOBID}: output and error logs for each job The file "deephyper-sorted.csv" contains the details of each HydraGNN model built and tested by HPO, obtained by merging the (*.csv) filed from each HPO run executed. Out of all the HPO trials, we selected 10 to continue the training of the respective HydraGNN models. Due to limited computational budget available in the LRN070 allocation we could not complete the training till convergence for all these 10 selected models. The folder "models" contains multiple sub-folders, one per each HydraGNN model trained. Each model sub-folder contains the parameters of each HydraGNN model, with multiple checkpoint-restarts. The list of sub-folders are as follows: 1) multidataset_hpo-BEST1-fp64 2) multidataset_hpo-BEST2-fp64 3) multidataset_hpo-BEST3-fp64 4) multidataset_hpo-BEST4-fp64 5) multidataset_hpo-BEST5-fp64 6) multidataset_hpo-BEST6-fp64 7) multidataset_hpo-BEST7-fp64 8) multidataset_hpo-BEST8-fp64 9) multidataset_hpo-BEST9-fp64 10) multidataset_hpo-BEST10-fp64 Within each one of these folders, additional auxiliary log files are provided with descriptions about how the training proceeded. The lead PaiNN-model is contained inside "multidataset_hpo-BEST6-fp64". The file "mlp_branch_weights" contains the parameters of the multi-layer perceptron (MLP) used to reconcile the predictions of the 16 output decoding heads of the HydragNN architectures. The MLP takes in input the chemical composition of the atomistic structure and predicts averaging weights to linearly mix the predictions of each output decoding head toward consolidating them into a single one. The folder "1.1billion-structure-inference" contains 1.1 billion atomistic structures randomly generated. Each structures is associated with energy and forces predicted with the lead-PaiNN model combined with the MLP model for reconciliation of the multi-branch predictions generated by the 16 output decoding heads. The folder "1.1billion-structure-inference" contains 9,300 (*.tar.gz) subdirectories, one per Frontier compute node used to execute the inference at exascale. Once uncompressed, each (*.tar.gz) subdirectory contains an ADIOS2 (*.bp) file container, where each atomistic structure is stored as a PyTorch-Geometric Data object. The file "export_dataset_environment_variables.sh" contains the environment variables that need to be set before running the HydraGNN code to reproduce the results provided in this dataset release. The code that can be used to load the ADIOS2 files, load HydraGNN models, and run inference is available at: https://github.com/ORNL/HydraGNN/releases/tag/v5.0

36 MATERIALS SCIENCE↗

Untangling the Galaxy. II. Structure within 3 kpc

We present the results of the hierarchical clustering analysis of the Gaia DR2 data to search for clusters, comoving groups, and other stellar structures. The current paper builds on the sample from the previous work, extending it in distance from 1 to 3 kpc and increasing the number of identified structures up to 8292. To aid in the analysis of the population properties, we developed a neural network called Auriga to robustly estimate the age, extinction, and distance of a stellar group based on the input photometry and parallaxes of the individual members. We apply Auriga to derive the properties of not only the structures found in this paper, but also previously identified open clusters. Through this work, we examine the temporal structure of the spiral arms. Specifically, we find that the Sagittarius Arm has moved by >500 pc in the last 100 Myr and the Perseus Arm has been experiencing a relative lull in star formation activity over the last 25 Myr. We confirm the findings of the previous paper on the transient nature of the spiral arms, with the timescale of transition of a few 100 Myr. Finally, we find a peculiar ~1 Gyr old stream of stars that appears to be heliocentric. Its origin is unclear.

79 ASTRONOMY AND ASTROPHYSICS↗

Simulation of Impedance Changes with Aging in Lithium Titanate-based Cells Using Physics-Based Dimensionless Modeling

Quantifying aging effects in lithium-ion cells with chemistries that have a flat open circuit potential is challenging. We implement a physics-based electrochemical model to track changes in the electrochemical impedance response of lithium titanate-based cells. Frequency domain equations of a pseudo two-dimensional model are made dimensionless, and the corresponding non-dimensional parameters are estimated using a Levenberg-Marquardt routine. The model weighs the relative contributions of changes in diffusion, ionic conduction within the electrolyte phase against solid phase electronic conduction towards cell aging. Solid-phase diffusion, charge transfer resistance and double layer capacitance at the solid-liquid interface are accounted for in the particle impedance. The estimation routine tracks dimensionless parameters using accelerated cycling data from full cells over 1000 cycles. The model can be deployed within a short time for state estimation using physics-based models without requiring prior knowledge of the battery chemistry, format, or capacity.

25 ENERGY STORAGE↗

RadSim: Math, Utility & RTK

This package includes three parts, (1) gov.llnl.math, (2)gov.llnl.utility and (3)gov.llnl.rtk, which are utilized in the development of Radiation Detector Simulator (RadSim) project. RadSim is being developed to provide the capability to: (1) simulate radiation source emissions, (2) interpolate results from radiation transport tools into a common format to prepare incident flux, and (3) model radiation detector response to rapidly produce synthetic radiation measurement templates. RadSim is targeted for open-source release, which will enable researchers and industry partners to model gamma-ray detectors response to simulated flux from the transport tool of their choice. The techniques and implementation will be entirely transparent, which will allow for improvements and boutique modifications by future researchers beyond the lifespan of this specific project. The first tool of the package, gov.llnl.math, includes classes and functions to define and perform basic math operations. Some of the example features available in the package include defining statistical distributions and performing algebra and matrix operations, all of which are already accessible on publicly available software packages such as MATLAB and ROOT. The second package gov.llnl.utility includes tools commonly used to enable optimization and readability of various data structures such as Java lists and external xml files. Lastly, the gov.llnl.rtk package includes classes and functions to implement methods commonly used in radiation physics, such as data structures to represent and characterize photon spectra and tools to apply well-defined and published methods to calibrate a given spectra.

Cheung, Hoi Sing↗

Changes in short- and medium-range order in metallic liquids during undercooling

It has been widely speculated that dominant motifs, such as short-range icosahedral order, can influence glass formation and the properties of glasses. Experimental data on both fragile and strong undercooled liquids demonstrate corresponding changes in their thermophysical properties consistent with increasing development of a network of interconnect motifs based on molecular dynamics. Describing these regions of local order, how they connect, and how they are related to property changes have been challenging issues, both computationally and experimentally. Yet the consensus is that metallic liquids develop interconnected medium-range order consisting of some regions with lower mobility with deeper undercooling. Less well understood is how these motifs (or “crystal genes”) in the liquid can inhibit nucleation in the deeply undercooled liquid or influence phase selection upon devitrification. These motifs tend to have local packing unlike stable compounds with icosahedral order tending to dominate the best glass formers. The underlying kinetic and thermodynamic forces that guide the formation of these motifs and how they interconnect during undercooling remain open questions.

36 MATERIALS SCIENCE↗

FREDA: A Web Application for the Processing, Analysis, and Visualization of Fourier‐Transform Mass Spectrometry Data

The high-resolution measurement capability of Fourier-transform mass spectrometry (FT-MS) has made it a necessity for exploring the molecular composition of complex organic mixtures, like soil, plant, aquatic, and petroleum samples. This demand has driven a need for informatics tools to explore and analyze FT-MS data in a robust and reproducible manner. FREDA is an interactive web application developed to enable spectrometrists to format, process, and explore their FT-MS data without the need for statistical programming expertise. FREDA was built to explore outputs from a molecular identification tool, like CoreMS, and provide a suite of methods to filter data, compute chemical properties of peaks, statistically compare samples and groups of samples, conduct exploratory data analysis, and download the results with a report detailing all steps conducted. To demonstrate the utility of FREDA, an example analysis was conducted using FT-MS data from a soil microbiology study of samples collected in two different soil depths at the Sphagnum bog forest north of Grand Rapids, Minnesota. Differences between the two depths are observed using Kendrick, Gibbs free energy, and van Krevelen plots. G-tests are used to quantify a significant difference between the groups. All analyses and plotting are conducted using only the FREDA application. FREDA is an open-source and readily available web application that allows users to explore and make statistically valid conclusions about their FT-MS data. The application is available online (https://map.emsl.pnnl.gov/app/freda) with a tutorial web series (https://youtu.be/k5HLE2kNSBY?si=yB6sGoyvzxrFf5MP) and freely accessible code on Github (https://github.com/EMSL-Computing/FREDA).

47 OTHER INSTRUMENTATION↗

Incommensurate charge-stripe correlations in the kagome superconductor CsV 3 Sb 5–x Sn x

The class of AV 3 Sb 5 (A=K, Rb, Cs) kagome metals hosts unconventional charge density wave states seemingly intertwined with their low temperature superconducting phases. The nature of the coupling between these two states and the potential presence of nearby, competing charge instabilities however remain open questions. This phenomenology is strikingly highlighted by the formation of two ‘domes’ in the superconducting transition temperature upon hole-doping CsV 3 Sb 5 . Here we track the evolution of charge correlations upon the suppression of long-range charge density wave order in the first dome and into the second of the hole-doped kagome superconductor CsV 3 Sb 5–x Sn x . Initially, hole-doping drives interlayer charge correlations to become short-ranged with their periodicity diminished along the interlayer direction. Beyond the peak of the first superconducting dome, the parent charge density wave state vanishes and incommensurate, quasi-1D charge correlations are stabilized in its place. These competing, unidirectional charge correlations demonstrate an inherent electronic rotational symmetry breaking in CsV 3 Sb 5 , and reveal a complex landscape of charge correlations within its electronic phase diagram. Our data suggest an inherent 2k ƒ charge instability and competing charge orders in the AV 3 Sb 5 class of kagome superconductors.

36 MATERIALS SCIENCE↗

Exploring electron-beam induced modifications of materials with machine-learning assisted high temporal resolution electron microscopy

Directed atomic fabrication using an aberration-corrected scanning transmission electron microscope (STEM) opens new pathways for atomic engineering of functional materials. In this approach, the electron beam is used to actively alter the atomic structure through electron beam induced irradiation processes. One of the impediments that has limited widespread use thus far has been the ability to understand the fundamental mechanisms of atomic transformation pathways at high spatiotemporal resolution. Here, we develop a workflow for obtaining and analyzing high-speed spiral scan STEM data, up to 100 fps, to track the atomic fabrication process during nanopore milling in monolayer MoS 2 . An automated feedback-controlled electron beam positioning system combined with deep convolution neural network (DCNN) was used to decipher fast but low signal-to-noise datasets and classify time-resolved atom positions and nature of their evolving atomic defect configurations. Through this automated decoding, the initial atomic disordering and reordering processes leading to nanopore formation was able to be studied across various timescales. Using these experimental workflows a greater degree of speed and information can be extracted from small datasets without compromising spatial resolution. This approach can be adapted to other 2D materials systems to gain further insights into the defect formation necessary to inform future automated fabrication techniques utilizing the STEM electron beam.

36 MATERIALS SCIENCE↗

The Coastal Carbon Library and Atlas: Open source soil data and tools supporting blue carbon research and policy

Abstract Quantifying carbon fluxes into and out of coastal soils is critical to meeting greenhouse gas reduction and coastal resiliency goals. Numerous ‘blue carbon’ studies have generated, or benefitted from, synthetic datasets. However, the community those efforts inspired does not have a centralized, standardized database of disaggregated data used to estimate carbon stocks and fluxes. In this paper, we describe a data structure designed to standardize data reporting, maximize reuse, and maintain a chain of credit from synthesis to original source. We introduce version 1.0.0. of the Coastal Carbon Library, a global database of 6723 soil profiles representing blue carbon‐storing systems including marshes, mangroves, tidal freshwater forests, and seagrasses. We also present the Coastal Carbon Atlas, an R‐shiny application that can be used to visualize, query, and download portions of the Coastal Carbon Library. The majority (4815) of entries in the database can be used for carbon stock assessments without the need for interpolating missing soil variables, 533 are available for estimating carbon burial rate, and 326 are useful for fitting dynamic soil formation models. Organic matter density significantly varied by habitat with tidal freshwater forests having the highest density, and seagrasses having the lowest. Future work could involve expansion of the synthesis to include more deep stock assessments, increasing the representation of data outside of the U.S., and increasing the amount of data available for mangroves and seagrasses, especially carbon burial rate data. We present proposed best practices for blue carbon data including an emphasis on disaggregation, data publication, dataset documentation, and use of standardized vocabulary and templates whenever appropriate. To conclude, the Coastal Carbon Library and Atlas serve as a general example of a grassroots F.A.I.R. (Findable, Accessible, Interoperable, and Reusable) data effort demonstrating how data producers can coordinate to develop tools relevant to policy and decision‐making.

Holmquist, James R.↗

IM3 Open Source Data Center Atlas

IM3 Open Source Data Center Atlas Description This dataset contains locations of existing data center facilities in the United States. Data center locations were derived from OpenStreetMap (OSM), a crowd-sourced database. Data points from OSM are processed in various ways to determine additional variables provided in the data including: facility area (square feet), associated US county, and US state. This dataset can be used to identify areas of concentrated data center development and inform government and private sector planning strategies for future buildout of data centers and the infrastructure necessary to support it. Usage Notes Validation of OSM-derived data center locations is an ongoing development under the IM3 project, and the database will be updated as new information becomes available. In some instances, both the data center area (e.g., campus) and individual data center buildings are included as overlapping areas in the database. Both values are retained. Data center points, buildings, and campus areas are provided as separate layers in the downloadable data package. Note that data items are not necessarily complete across layers. That is, a specific data center may only be present as a single point geometry in the "point" layer while other data centers are represented in both the campus and building layers. In some cases, data center campuses and/or buildings straddle a county boundary line. Mappings to both counties are retained in the database as separate rows. These data rows will have the same data center id information, but each will have different county information. Crowd-sourced data, by nature, relies on individuals and communities to provide information. As a result, some data may be missing where it has not yet been reported. As we collect information on additional data center locations and as OSM receives additional contributions, the database will be updated to capture additional data points not yet shown. Data items will occasionally be removed from OSM if they are misidentified, if they no longer exist, if they are duplicates of another item, or similar. For that reason, updated versions of this database may not contain all data center locations included in previous versions. Technical Information Data is available for download under the following formats: GeoPackage (GPKG) CSV Geospatial data is provided in the WGS84 (EPSG:4326) coordinate reference system. The GeoPackage download contains the following layers. See usage notes for more information. "point" "building" "campus" The "point" layer includes all data from OSM that had POINT geometry type (i.e., individual coordinates). The "building" layer includes all OSM data that did not have POINT geometry and where the building tag in the OSM export was neither equal to "no" or null. Data that did not meet the "point" or "building" qualification was assumed to be a facility campus and included in the "campus" layer. The dataset contains the following parameters. Variables provided by OSM are labeled with (OSM-provided). id - unique identification number (OSM-provided with prefix of "node/", "relation/" and similar attributes removed) state - name of US state state_abb - two letter US state abbreviation state_id - state ID number county - name of US county county_id - county ID number ref - reference numbers or codes (OSM-provided) operator - the name of the company, corporation, or person in charge facility (OSM-provided) name - name of facility (OSM-provided) sqft - surface area of facility polygon, measured in square feet. Only available for "building" and "campus" layers lat - latitude of data centroid point lon - longitude of data centroid point type – represented spatial information. One of "point", "building", or "campus". geometry – POLYGON geometry of area footprint (in "campus" and "building" layers) or POINT geometry of locations (in "point" layer). This parameter is not included in the csv download. Attribution Data center locations were derived from OpenStreetMap, which is made available at openstreetmap.org under the Open Database License (ODbL). US state and county boundary information was collected from the US Census Bureau for the year 2024, which is made publicly available at https://www.census.gov/geographies/mapping-files.html Acknowledgment IM3 is a multi-institutional effort led by Pacific Northwest National Laboratory and supported by the U.S. Department of Energy's Office of Science as part of research in MultiSector Dynamics, Earth and Environmental Systems Modeling Program. License The IM3 Open Source Data Center Atlas is made available under the Open Database License: http://opendatacommons.org/licenses/odbl/1.0/. Disclaimer This material was prepared as an account of work sponsored by an agency of the United States Government. Neither the United States Government nor the United States Department of Energy, nor the Contractor, nor any or their employees, nor any jurisdiction or organization that has cooperated in the development of these materials, makes any warranty, express or implied, or assumes any legal liability or responsibility for the accuracy, completeness, or usefulness or any information, apparatus, product, software, or process disclosed, or represents that its use would not infringe privately owned rights. Reference herein to any specific commercial product, process, or service by trade name, trademark, manufacturer, or otherwise does not necessarily constitute or imply its endorsement, recommendation, or favoring by the United States Government or any agency thereof, or Battelle Memorial Institute. The views and opinions of authors expressed herein do not necessarily state or reflect those of the United States Government or any agency thereof. PACIFIC NORTHWEST NATIONAL LABORATORYoperated byBATTELLEfor theUNITED STATES DEPARTMENT OF ENERGYunder Contract DE-AC05-76RL01830

Mongird, Kendall [Pacific Northwest National Labor↗

Preliminary report on applications of machine learning techniques to the Nevada play fairway analysis

We are applying machine learning (ML) techniques, including training set augmentation and artificial neural networks, to mitigate key challenges in the Nevada play fairway project. The study area includes ~85 active geothermal systems as potential training sites and >12 geologic, geophysical, and geochemical features. The main goal is to develop an algorithmic approach to identify new geothermal systems in the Great Basin region. Major objectives include: 1) integrate ML techniques into the geothermal community; 2) develop open community datasets, whereby all play fairway and ML datasets and algorithms are publicly released and available for modification by various user groups; 3) identify data acquisition targets with high value for future work; 4) identify new signatures to detect blind geothermal systems; and 5) foster new capabilities for characterizing subsurface temperature and permeability. Initially, ML techniques are being applied to the same play fairway datasets and workflow. ML will then be applied to both enhanced and additional datasets, with modification of the PFA workflow to incorporate the new datasets. Finally, ML will be applied to define new workflows using the enhanced and additional datasets. An algorithmic approach that empirically learns to estimate weights of influence for diverse parameters can potentially scale and perform better than the play fairway analysis. Initial work on this project has involved 1) evaluating potential positive and negative training sites, 2) transformation of datasets into formats suitable for ML, and 3) initial development and testing of ML techniques.

58 GEOSCIENCES↗

UCB-GLOBES: An open-access mass spectral database of identified and unidentified atmospheric organic compounds

Chemical characterization of atmospheric organic aerosols using gas chromatography with 70 eV electron ionization mass spectrometry (GC/EI-MS) has been used for decades in advancing molecular marker detection and identification, though primarily through suspect screening and/or targeted analyses. To advance non-targeted analyses of environmental samples, we have catalogued approximately 27 000 mass spectra (MS) of the trimethylsilyl derivatives of semi-volatile organic aerosol (OA) analytes in the open-access University of California Berkeley Goldstein Library of Organic Biogenic Environmental Spectra (UCB-GLOBES). Analytes were observed in ambient samples from the U.S. and the Central Amazon and/or laboratory simulations of secondary OA (SOA) formation. These samples are representative of OA under urban and biomass burning influences as well as SOA derived from biogenic precursors (e.g., isoprene, monoterpenes, sesquiterpenes) and biomass burning intermediates. MS are documented in UCB-GLOBES without regard to known chemical identity, annotated with extensive metadata such as sample source/experimental conditions, any structural information gained from MS analyses, and predicted chemical properties such as average carbon oxidation state and carbon number. UCB-GLOBES MS are compatible for importing into the NIST MS Search program, and we have also provided a Jupyter Notebook for MS visualization and comparisons. We demonstrate the utility of UCB-GLOBES through MS reanalyses of prior analytes observed in ambient data, finding a 20 % reduction in the number of analytes assigned to OA source categories reliant solely on time series correlation and an overall 11 % increase in new MS-based OA source categorization for the Southeast U.S. For 1513 analytes observed previously in the Central Amazon, we found 375 MS matches using UCB-GLOBES vs. 136 MS matches during prior analyses, representing a 14 % gain in newly confirmed or newly categorized OA species. While OA from laboratory oxidation experiments in UCB-GLOBES are highly diverse chemically, on average only 29 % of UCB-GLOBES MS have a mass spectral match to another MS entry in UCB-GLOBES and/or in databases of known compounds (i.e. NIST MS Database, Adams Essential Oil, MANE Flavor and Fragrance Company). This indicates that roughly 70 % of UCB-GLOBES MS are unique thus far, not observed more than once among the laboratory oxidation samples and ambient data in UCB-GLOBES MS. Further, only 18 % can be positively identified using these databases or known authentic standards. This points to a large gap between these laboratory simulations and ambient OA. Overall, the UCB-GLOBES database can be utilized for improving confidence in OA source categorization and/or identification, novel chemical marker discovery, tracking chemical diversity, de novo structure and properties prediction, and improving MS search and matching algorithms. This can ultimately inform future research priorities for the chemical characterization of atmospheric organic samples.

Mass spectrometry↗

A matched-filter approach to radio variability and transients: searching for orphan afterglows in the VAST Pilot Survey

Radio transient searches using traditional variability metrics struggle to recover sources whose evolution time-scale is significantly longer than the survey cadence. Motivated by the recent observations of slowly evolving radio afterglows at gigahertz frequency, we present the results of a search for radio variables and transients using an alternative matched-filter approach. We designed our matched-filter to recover sources with radio light curves that have a high-significance fit to power-law and smoothly broken power-law functions; light curves following these functions are characteristic of synchrotron transients, including ‘orphan’ gamma-ray burst afterglows, which were the primary targets of our search. Applying this matched-filter approach to data from Variables and Slow Transients Pilot Survey conducted using the Australian SKA Pathfinder, we produced five candidates in our search. Subsequent Australia Telescope Compact Array observations and analysis revealed that: one is likely a synchrotron transient; one is likely a flaring active galactic nucleus, exhibiting a flat-to-steep spectral transition over 4 months; one is associated with a starburst galaxy, with the radio emission originating from either star formation or an underlying slowly evolving transient; and the remaining two are likely extrinsic variables caused by interstellar scintillation. The synchrotron transient, VAST J175036.1–181454, has a multifrequency light curve, peak spectral luminosity, and volumetric rate that is consistent with both an off-axis afterglow and an off-axis tidal disruption event; interpreted as an off-axis afterglow would imply an average inverse beaming factor $\langle f^{-1}_{\text{b}} \rangle = 860^{+1980}_{-710}$, or equivalently, an average jet opening angle of $\langle \theta _{\textrm {j}} \rangle = 3^{+4}_{-1}\,$ deg.

79 ASTRONOMY AND ASTROPHYSICS↗

Modeling a ring magnet in ALEGRA

We show here that Sandia's ALEGRA software can be used to model a permanent magnet in 2D and 3D, with accuracy matching that of the open-source commercial software FEMM. This is done by conducting simulations and experimental measurements for a commercial-grade N42 neodymium alloy ring magnet with a measured magnetic field strength of approximately 0.4 T in its immediate vicinity. Transient simulations using ALEGRA and static simulations using FEMM are conducted. Comparisons are made between simulations and measurements, and amongst the simulations, for sample locations in the steady-state magnetic field. The comparisons show that all models capture the data to within 7%. The FEMM and ALEGRA results agree to within approximately 2%. The most accurate solutions in ALEGRA are obtained using quadrilateral or hexahedral elements. In the case where iron shielding disks are included in the magnetized space, ALEGRA simulations are considerably more expensive because of the increased magnetic diffusion time, but FEMM and ALEGRA results are still in agreement. The magnetic field data are portable to other software interfaces using the Exodus file format.

36 MATERIALS SCIENCE↗

In Situ Ultra-Small- and Small-Angle X-ray Scattering Study of ZnO Nanoparticle Formation and Growth through Chemical Bath Deposition in the Presence of Polyvinylpyrrolidone

ZnO inverse opals combine the outstanding properties of the semiconductor ZnO with the high surface area of the open-porous framework, making them valuable photonic and catalysis support materials. One route to produce inverse opals is to mineralize the voids of close-packed polymer nanoparticle templates by chemical bath deposition (CBD) using a ZnO precursor solution, followed by template removal. To ensure synthesis control, the formation and growth of ZnO nanoparticles in a precursor solution containing the organic additive polyvinylpyrrolidone (PVP) was investigated by in situ ultra-small- and small-angle X-ray scattering (USAXS/SAXS). Before that, we studied the precursor solution by in-house SAXS at T = 25 °C, revealing the presence of a PVP network with semiflexible chain behavior. Heating the precursor solution to 58 °C or 63 °C initiates the formation of small ZnO nanoparticles that cluster together, as shown by complementary transmission electron microscopy images (TEM) taken after synthesis. The underlying kinetics of this process could be deciphered by quantitatively analyzing the USAXS/SAXS data considering the scattering contributions of particles, clusters, and the PVP network. A nearly quantitative description of both the nucleation and growth period could be achieved using the two-step Finke–Watzky model with slow, continuous nucleation followed by autocatalytic growth.

36 MATERIALS SCIENCE↗

Data from: “Enabling FAIR data in Earth and environmental science with community-centric (meta)data reporting formats”

This dataset contains supplementary information for a manuscript describing the ESS-DIVE (Environmental Systems Science Data Infrastructure for a Virtual Ecosystem) data repository's community data and metadata reporting formats. The purpose of creating the ESS-DIVE reporting formats was to provide guidelines for formatting some of the diverse data types that can be found in the ESS-DIVE repository. The 6 teams of community partners who developed the reporting formats included scientists and engineers from across the Department of Energy National Lab network. Additionally, during the development process, 247 individuals representing 128 institutions provided input on the formats. The primary files in this dataset are 10 data and metadata crosswalk for ESS-DIVE’s reporting formats (all files ending in _crosswalk.csv). The crosswalks compare elements used in each of the reporting formats to other related standards and data resources (e.g., repositories, datasets, data systems). This dataset also contains additional files recommended by ESS-DIVE’s file-level metadata reporting format. Each data file has an associated dictionary (files ending in _dd.csv) which provide a brief description of each standard or data resource consulted in the data reporting format development process. The flmd.csv file describes each file contained within the dataset.

54 ENVIRONMENTAL SCIENCES↗

Survey of Surveys: I. The largest compilation of radial velocities for the Galaxy

In the present-day panorama of large spectroscopic surveys, the amount, diversity, and complexity of the available data continuously increase. The overarching goal of studying the formation and evolution of our Galaxy is hampered by the heterogeneity of instruments, selection functions, analysis methods, and measured quantities. Here, we present a comprehensive catalogue, the Survey of Surveys (SoS), built by homogeneously merging the radial velocity (RV) determinations of the largest ground-based spectroscopic surveys to date, such as APOGEE, GALAH, Gaia-ESO, RAVE, and LAMOST, using Gaia as a reference. This pilot study serves to prove the concept and to test the methodology that we plan to apply in the future to the stellar parameters and abundance ratios as well. We have devised a multi-staged procedure that includes: (i) the cross match between Gaia and the spectroscopic surveys using the official Gaia cross-match algorithm, (ii) the normalisation of uncertainties using repeated measurements or the three-cornered hat method, (iii) the cross calibration of the RVs as a function of the main parameters on which depend (magnitude, effective temperature, surface gravity, metallicity, and signal-to-noise ratio) to remove trends and zero point offsets, and (iv) the comparison with external high-resolution samples, such as the Gaia RV standards and the Geneva-Copenhagen survey, to validate the homogenisation procedure and to calibrate the RV zero-point of the SoS catalogue. We provide the largest homogenised RV catalogue to date, containing almost 11 million stars, of which about half come exclusively from Gaia and half in combination with the ground-based surveys. We estimate the accuracy of the RV zero-point to be about 0.16–0.31 km s –1 and the RV precision to be in the range 0.05–1.50 km s –1 depending on the type of star and on its survey provenance. We validate the SoS RVs with open clusters from a high resolution homogeneous samples and provide the systemic velocity of 55 individual open clusters. Additionally, we provide median RVs for 532 clusters recently discovered by Gaia data. The SoS is publicly available and ready to be applied to various research projects, such as the study of star clusters, Galactic archaeology, stellar streams, or the characterisation of planet-hosting stars, to name a few. We also plan to include survey updates and more data sources in future versions of the SoS.

79 ASTRONOMY AND ASTROPHYSICS↗

Photometric and Kinematic Study of the Open Clusters SAI 44 and SAI 45

We carry out a detailed photometric and kinematic study of the poorly studied sparse open clusters SAI 44 and SAI 45 using ground-based BVR {sub c} I {sub c} data supplemented by archival data from Gaia eDR3 and Pan-STARRS. The stellar memberships are determined using a statistical method based on Gaia eDR3 kinematic data, and we found 204 members in SAI 44 while only 74 members are identified in SAI 45. The average distances to SAI 44 and SAI 45 are calculated to be 3670 ± 184 and 1668 ± 47 pc. The logarithmic age of the clusters are determined to be 8.82 ± 0.10 and 9.07 ± 0.10 yr for SAI 44 and SAI 45, respectively. The color–magnitude diagram of SAI 45 hosts an extended main-sequence turnoff (eMSTO). The apparent age spread is found to be similar to the apparent age spread predicted on the basis of the age spread and cluster age relation predicted by rotation models. This indicates that eMSTO is a stellar evolution rather than star formation phenomenon in SAI 45. We conclude that eMSTO in SAI 45 is mainly caused by the different rotation rates of stars as the SYCLIST synthetic population with different rotation rates was able to reproduce the observed eMSTO, and stars in the red part of the eMSTO were preferentially concentrated in the inner region, which again hints at different rotations being the reason for the extension in the upper MS. This finding supports the theory attributing the origin of eMSTO to the different rotations of eMSTO stars. The mass function slopes are obtained as -2.24 ± 0.66 and -2.58 ± 3.20 in the mass rages 2.426–0.990 M{sub ⊙} and 2.167–1.202 M{sub ⊙} for SAI 44 and SAI 45, respectively. SAI 44 exhibits the signature of mass segregation while we found weak evidence of mass segregation in SAI 45 possibly due to tidal stripping. The dynamical relaxation times of these clusters indicate that both clusters are in a dynamically relaxed state. Using the AD-diagram method, the apex coordinates are found to be ([[Formula]], -[[Formula]]) for SAI 44 and (-[[Formula]], -[[Formula]]) for SAI 45. The average space velocity components of the clusters SAI 44 and SAI 45 are calculated in units of km s{sup -1} as (-15.14 ± 3.90, -19.43 ± 4.41, -20.85 ± 4.57) and (28.13 ± 5.30, -9.78 ± 3.13, -19.59 ± 4.43), respectively.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗