Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Data structures”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

LinkML: an open data modeling framework

Background Scientific research relies on well-structured, standardized data; however, much of it is stored in formats such as free-text lab notebooks, nonstandardized spreadsheets, or data repositories. This lack of structure challenges interoperability, making data integration, validation, and reuse difficult. Findings LinkML (Linked Data Modeling Language) is an open framework that simplifies the process of authoring, validating, and sharing data. LinkML can describe a range of data structures, from flat, list-based models to complex, interrelated, and normalized models that utilize polymorphism and compound inheritance. It offers an approachable syntax that is not tied to any one technical architecture and can be integrated seamlessly with many existing frameworks. The LinkML syntax provides a standard way to describe schemas, classes, and relationships, allowing modelers to build well-defined, stable, and optionally ontology-aligned data structures. Once defined, LinkML schemas may be imported into other LinkML schemas. These key features make LinkML an accessible platform for interdisciplinary collaboration and a reliable way to define and share data semantics. Conclusions LinkML helps reduce heterogeneity, complexity, and the proliferation of single-use data models while simultaneously enabling compliance with FAIR (Findable, Accessible, Interoperable, and Reusable) data standards. LinkML has seen increasing adoption in various fields, including biology, chemistry, biomedicine, microbiome research, finance, electrical engineering, transportation, and commercial software development. In short, LinkML makes implicit models explicitly computable and allows data to be standardized at their origin. LinkML documentation and code are available at https://linkml.io/.

AI-ready data↗

An Active Learning-Based Streaming Pipeline for Reduced Data Training of Structure Finding Models in Neutron Diffractometry

Structure determination workloads in neutron diffractometry are computationally expensive and routinely require several hours to many days to determine the structure of a material from its neutron diffraction patterns. The potential for machine learning models trained on simulated neutron scattering patterns to significantly speed up these tasks have been reported recently. However, the amount of simulated data needed to train these models grows exponentially with the number of structural parameters to be predicted and poses a significant computational challenge. To overcome this challenge, we introduce a novel batch-mode active learning (AL) policy that uses uncertainty sampling to simulate training data drawn from a probability distribution that prefers labelled examples about which the model is least certain. We confirm its efficacy in training the same models with ∼ 75% less training data while improving the accuracy. We then discuss the design of an efficient stream-based training workflow that uses this AL policy and present a performance study on two heterogeneous platforms to demonstrate that, compared with a conventional training workflow, the streaming workflow delivers ∼ 20% shorter training time without any loss of accuracy.

Wang, Tianle [Brookhaven National Laboratory (BNL)↗

Massive compression for high data rate macromolecular crystallography (HDRMX): impact on diffraction data and subsequent structural analysis

New higher-count-rate, integrating, large-area X-ray detectors with framing rates as high as 17400 images per second are beginning to be available. These will soon be used for specialized macromolecular crystallography experiments but will require optimal lossy compression algorithms to enable systems to keep up with data throughput. Some information may be lost. Can we minimize this loss with acceptable impact on structural information? To explore this question, we have considered several approaches: summing short sequences of images, binning to create the effect of larger pixels, use of JPEG-2000 lossy wavelet-based compression, and use of Hcompress, which is a Haar-wavelet-based lossy compression borrowed from astronomy. We also explore the effect of the combination of summing, binning, and Hcompress or JPEG-2000. In each of these last two methods one can specify approximately how much one wants the result to be compressed from the starting file size. These provide particularly effective lossy compressions that retain essential information for structure solution from Bragg reflections.

47 OTHER INSTRUMENTATION↗

PPI DataHub Project Data Package: High-density Lipoprotein (HDL) Structure and Function Proteomics

The purpose of this experiment was to investigate how the interactions between APOA1 and APOA2 on the surface of high-density lipoproteins (HDL) impact particle function. Interactions were investigated on HDL isolated from human blood plasma using structural proteomics tools such as chemical cross-linking and limited proteolysis (LiP). The structural proteomics data was acquired using a Q-Exactive HF-X mass spectrometer and data was processed and compiled using MaxQuant sofware (v.1.6.17.0). Processed datasets are openly accessible from the download button (~2.8 GB) and contain secondary processed LiP and global proteomic results files and supporting metadata materials. Processed data downloads include a sample naming key, processed MaxQuant results/parameters, and protein annotated relative abundance files.

59 BASIC BIOLOGICAL SCIENCES↗

Physics-Informed Gaussian Process Inference of Liquid Structure from Scattering Data

We present a nonparametric Bayesian framework to infer radial distribution functions from experimental scattering measurements with uncertainty quantification using nonstationary Gaussian processes. The Gaussian process prior mean and kernel functions are designed to mitigate well-known numerical challenges with the Fourier transform, including discrete measurement binning and detector windowing, while encoding fundamental yet minimal physical knowledge of the liquid structure. We demonstrate uncertainty propagation of the Gaussian process posterior to unmeasured quantities of interest. Experimental radial distribution functions of liquid argon and water with uncertainty quantification are provided as both a proof of principle for the method and a benchmark for molecular models.

Chemical structure↗

Nuclear Structure and Decay Data for A=169 Isobars

Experimental data pertaining to all nuclei with mass number A=169 (Eu, Gd, Tb, Dy, Ho, Er, Tm, Yb, Lu, Hf, Ta, W, Re, Os, Ir, Pt) have been evaluated. Level schemes from both radioactive decay and reaction studies are presented, along with associated tables of experimental data and adopted properties for levels and γ rays. The present evaluation for A=169 supersedes the 2008 evaluation, 2008Ba31, by C.M. Baglin. A few highlights of this evaluation: More extensive work on ε decay from 169W is needed and new experimental work will be required to resolve a discrepancy between the J π values deduced for a 180-keV level in 169Ta based on extensive band structure from (HI,xnγ) work (J π =1/2−) and TDPAD measurements (J=5/2). Low lying states of 169Os were studied via fine structure of 173Pt α decay in 2014ThZZ. The Eαs feeding the g.s. of 169Os in 2008Ba31 are separated well into two consistent groups to feed the g.s. and the newly proposed state at 34.84 keV. Based on the studies of 2014ThZZ and 2021Zh52, the g.s. spin-parity assignment of 169Os has been proposed to be (7/2−) from (5/2−). The 169Ir g.s. half-life and alpha emission branching reported in 2012Th13 from 173Au α decay measurements are preferred over the values in 2005Sc22. The reported half-life value in 2005Sc22 for 169Ir g.s. is discrepant and the research work was carried out in the same lab of 2012Th13.

Basunia, M Shamsuzzoha↗

Nuclear Structure and Decay Data for A=240

The experimental reaction and decay studies producing nuclei in the A=240 mass chain have been reviewed. Data on elements from uranium (Z=92) to einsteinium (Z=99) are included, and level and decay schemes are presented for these nuclides. Furthermore, this work supersedes the previous evaluation for this mass chain (2008Si25).

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Nuclear Structure and Decay Data for A=216 Isobars

Experimental data from reaction and decay studies on nuclei with A=216 have been reviewed. Elements included in this review span from mercury (Z=80) to uranium (Z=92). Based on the published data, level and decay schemes are presented for the evaluated nuclides. In conclusion, this work updates and supersedes the previous A=216 evaluation (2007Wu02).

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Nuclear Structure and Decay Data for A=35 Isobars

Here, this work presents a comprehensive and critical evaluation of experimental nuclear spectroscopic data from reactions and decays for all 11 known nuclides with mass number 35 (Ne, Na, Mg, Al, Si, P, S, Cl, Ar, K, Ca). Recommended values are produced for level energies, spins and parities, half-lives, and radiation properties including energies, branching ratios, and multipolarities of γ rays, as well as characteristics of β radiation decays, based on a rigorous assessment of all available experimental data. Discrepancies among existing results are carefully addressed. This work supersedes earlier full evaluations of A=35 published by 2011Ch48, 1990En08 (also 1998En04 update) and 1978En02.

Sun, Lijie [Michigan State University, East Lansin↗

Nuclear Structure and Decay Data for A=220 Isobars

Here, this work presents a comprehensive and critical evaluation of experimental nuclear spectroscopic data from reactions and decays for all 11 known nuclides with mass number A=220 (Pb, Bi, Po, At, Rn, Fr, Ra, Ac, Th, Pa, and Np). Recommended values are produced for level energies, spins and parities, half-lives, and radiation properties including energies, branching ratios, and multipolarities of γ rays, as well as characteristics of β and α radiation decays, based on all available experimental data. This work supersedes previous A=220 evaluations: 2011Br05, 1997Ar04, 1986Ma45, 1976El04.

Chen, J. [Michigan State University, East Lansing,↗

Nuclear Structure and Decay Data for A=230

The experimental reaction and decay studies producing nuclei in the A=230 mass chain have been reviewed. Data on elements from radon (Z=86) to americium (Z=95) are included, and level and decay schemes are presented for these nuclides. This work supersedes the previous evaluation for this mass chain (2012Br12).

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

An Autonomous MCP Bridge to Rucio: Enhancing Data Management Accessibility for High Energy Physics

The Rucio Data Management System [1] is an important tool used by High Energy Physics experiments, including those at Fermi National Accelerator Laboratory, to store and manage exabyte-scale scientific datasets. Despite its central role in coordinating data across globally distributed storage sites, Rucio's command line interface (CLI) presents a steep learning curve, and makes it difficult for scientists to navigate through. To solve this issue, a containerized Model Context Protocol (MCP) [2] server was built that connects Large Language Models directly to Rucio, allowing AI agents to handle data tasks by using simple, natural language rather than memorized terminal commands. The core engineering focus of this project was moving the server away from slow terminal commands that require text parsing and replacing them with a native Python Client API toolset and a planned REST API framework. Moving to the Python API handles data operations directly in memory, which helps clear up formatting errors, provides the AI with clean, structured JSON data and speeds up tool execution. To prove that the system actually works, a benchmarking pipeline was also built with various questions to test the AI across four different model configurations. The questions included finding data scopes, tracking down specific datasets, and checking replication rules. Through benchmarking, early runs showed that with raw terminal text, the model would get confused and stuck, whereas switching to the Python API to feed the AI clean, structured data yielded massive improvement. By creating an intelligent and autonomous bridge to a storage network, this project shows how AI can be implemented in scientific data management, which ultimately helps scientists at Fermilab spend less time sorting through data and more time focusing on their experiments and analysis.

Akella, Kashyap [William Rainey Harper Coll.]↗

Vertical Structure of Local Disk Galaxies Revealed by DESI Imaging Data

The vertical structure of galactic disks is an important probe of the disk assembly history. Here we investigate the vertical structure of a sample of 79 local disk galaxies within 50 Mpc using data from the Dark Energy Spectroscopic Instrument Legacy Imaging Surveys. Vertical luminosity profiles as a function of radius in the g, r, and z bands are extracted and fitted using a single-component sech2 model to determine scale height and its radial variation (i.e., flaring). Our measurements indicate local galactic disks are overall thin with negligible flaring. The median scale heights at 1 R e are 0.21, 0.22, and 0.22 kpc in the g, r, and z bands, respectively, while the median radial gradients of scale height are −0.006, 0.003, and 0.001 in these bands. These values are consistent with that of the geometric thin disk of the Milky Way represented by metal-rich, low-[α/Fe] populations, confirming the weak flaring of the geometric thin disk. A clear positive correlation of scale height with stellar mass is observed down to the low-mass end of 10 7 M ⊙ . These results provide a homogeneous benchmark for the vertical structure of nearby disk galaxies and establish the cosmological representativeness of the Milky Way thin disk.

79 ASTRONOMY AND ASTROPHYSICS↗

Multidimensional scaling informed by F -statistic: Visualizing grouped microbiome data with inference

Multidimensional scaling (MDS) is a widely used dimensionality reduction technique in microbial ecology data analysis that captures the multivariate structure of the data while preserving pairwise distances between samples. While improvements in MDS have enhanced the ability to reveal group-specific data patterns, these MDS-based methods require prior assumptions for inference, limiting their application in general microbiome analysis. Here, in this study, we introduce a new MDS-based ordination method, “F-informed MDS,” which configures the data distribution based on the F-statistic, the ratio of dispersion between groups sharing common and different characteristics. Using semisynthetic datasets, we demonstrate that the proposed method is robust to hyperparameter selection while maintaining statistical significance throughout the ordination process. Various quality metrics for evaluating dimensionality reduction confirm that F-informed MDS is comparable to state-of-the-art methods in preserving both local and global data structures. Its application to a diatom-associated bacterial community suggests the role of this new method in interpreting the community’s response to the host. Our approach offers a well-founded refinement of MDS that aligns with statistical test results, which can be beneficial for broader multidimensional data analyses in microbiology and ecology. This new visualization tool can be incorporated into standard microbiome data analyses.

Biological and medical sciences↗

Updated resources for exploring experimentally-determined PDB structures and Computed Structure Models at the RCSB Protein Data Bank

The Research Collaboratory for Structural Bioinformatics Protein Data Bank (RCSB PDB, RCSB.org), the US Worldwide Protein Data Bank (wwPDB, wwPDB.org) data center for the global PDB archive, provides access to the PDB data via its RCSB.org research-focused web portal. We report substantial additions to the tools and visualization features available at RCSB.org, which now delivers more than 227000 experimentally determined atomic-level three-dimensional (3D) biostructures stored in the global PDB archive alongside more than 1 million Computed Structure Models (CSMs) of proteins (including models for human, model organisms, select human pathogens, crop plants and organisms important for addressing climate change). In addition to providing support for 3D structure motif searches with user-provided coordinates, new features highlighted herein include query results organized by redundancy-reduced Groups and summary pages that facilitate exploration of groups of similar proteins. Newly released programmatic tools are also described, as are enhanced training opportunities.

Burley, Stephen K.↗

Machine learning approaches for crystallographic classification from synthetic 2D X-ray diffraction data

Crystallographic structure identification is crucial for understanding material properties; however, current methodologies often depend on labor-intensive and time-consuming analyses of 2D X-ray diffraction (XRD) patterns. To address these limitations, this study employs synthetic 2D XRD patterns combined with deep learning (DL) techniques to enable automated and high-throughput classification of the seven crystal systems and 230 space groups. We introduce the novel Auto Diffraction Pipeline, designed to generate synthetic 2D XRD spot patterns from crystallographic information files under diverse conditions, including varying zone axes, atomic substitution, atomic depletion and mechanical loading. These conditions enhance the realism of synthetic data, mitigating the scarcity of experimental datasets and enabling the creation of large representative training sets. Convolutional neural networks were trained and validated on these synthetic datasets to classify crystallographic structures across multiple scenarios. Our results demonstrate that integrating synthetic 2D XRD patterns with DL facilitates rapid, accurate and automated crystallographic classification, promoting the wider adoption of data-driven approaches in materials science.

Shahnazari, Ayoub [Univ. of Rochester, NY (United ↗