Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “science data analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Soil and groundwater environmental sensor data, Wax Lake Delta, Louisiana, March 2023 - March 2024

This study evaluates how environmental parameters that integrate biogeochemical processes vary with water table fluctuations in the freshwater Wax Lake Delta (WLD) in Louisiana, U.S.A. This data package contains seven *.csv files and one Excel file that compiles all the data from the individual .csv files. This dataset reports high frequency (15-min) observations of water level, soil redox potential, specific conductance, and pH made for one year along elevation transects located on the older, proximal (OT) and younger, distal (YT) ends of a deltaic island. Water depth relative to the ground surface (cm; HOBO U20L-04; error ± 0.4 cm), water pH and temperature (HOBO MX2501), and specific conductance and temperature (HOBO U24-001) sensors were installed in March 2023. Water depth was corrected for barometric pressure recorded by a separate logger secured to a platform above the highest water level. Soil redox probes (SWAP ORP-40-4-B) were also installed in March 2023. Each probe had four Pt sensors (2 mm width) placed at 10 cm, 20 cm, 30 cm, and 40 cm below the ground surface. Redox data were referenced to an external Ag0/AgCl (3M KCl) reference probe placed in saturated ground and recorded on CR1000X dataloggers (Campbell Scientific) powered by solar panels. A second reference probe was positioned near the primary reference probe for backup and data correction. The tops of the soil redox probes and soil moisture probes were flush with the soil surface so that sensors are reported at their indicated depths below ground surface. Here, we report data collected between 15 March 2023 to 15 March 2024 for all sensors, with some differences due to exact dates of sensor placement or data gaps associated with sensor malfunction. For example, water depth at OT4 was not recorded between March to November 2023. Data flags indicate whether a value is valid (1) or was excluded from data analysis in the associated manuscript (-1).

EARTH SCIENCE > LAND SURFACE > SOILS↗

Explainable Machine Learning for Functional Data

Black-box machine learning models are recognized as useful tools for prediction applications, but the algorithmic complexity of some models causes interpretation challenges. Explainability methods have been proposed to provide insight into these models, but there is little research focused on supervised modeling with functional data inputs. We argue that, especially in applications of high consequence, it is important to explicitly model the functional dependence in a black-box analysis to not obscure or misrepresent patterns in explanations. As such, we propose the V ariable importance E xplainable E lastic S hape A nalysis (VEESA) pipeline for training supervised machine learning models with functional inputs. The pipeline is an analysis process that includes the data preprocessing, modeling, and post-hoc explanations. The preprocessing is done using elastic functional principal components analysis, which accounts for vertical and horizontal variability in functional data and, ultimately, allows for explanations in the original data space that identify the important functional variability without bias due to correlated variables. Here, we demonstrate the pipeline on two high-consequence applications: explosives classification for national security and inkjet printer identification in forensic science. The applications exhibit the VEESA pipeline’s ability to provide an understanding of the characteristics of the functional data useful for prediction. Code for implementing the pipeline is available in the veesa R package (and supplemental python code).

Elastic Shape Analysis↗

From Chaos to Clarity: Autonomous Materials Discovery for Extreme Environments

The pursuit of advanced functional materials for energy applications demands an understanding of their behavior under the most challenging conditions. Extreme environments, characterized by intense radiation, high temperatures, and corrosive chemistries, push materials to their limits, often revealing unexpected behaviors and degradation pathways. Traditional materials research approaches, relying on trial-and-error experimentation, are often slow and resource-intensive, ill-suited to the complexities of extreme environments. This talk will explore the transformative potential of autonomous materials science in revolutionizing our understanding of materials synthesis and degradation in extreme environments. By integrating advanced microscopy techniques, artificial intelligence, and robotic experimentation, we can accelerate the discovery and design of resilient materials for a sustainable future. The presentation will highlight recent breakthroughs in autonomous microscopy, computer vision, and machine learning, showcasing their ability to unravel complex material transformations at the atomic scale. The talk will also delve into the challenges and opportunities associated with deploying autonomous systems to probe extreme environments, emphasizing the importance of robust algorithms, real-time data analysis, and adaptive experimentation. Our ultimate goal is to empower scientists with unprecedented capabilities to explore, understand, and engineer materials that can withstand the harshest conditions, paving the way for innovations in energy, aerospace, and beyond.

artificial intelligence↗

Generic and ML Workloads in an HPC Datacenter: Node Energy, Job Failures, and Node-Job Analysis

HPC datacenters offer a backbone to the modern digital society. Increasingly, they run Machine Learning (ML) jobs next to generic, compute-intensive workloads, supporting science, business, and other decision-making processes. However, understanding how ML jobs impact the operation of HPC datacenters, relative to generic jobs, remains desirable but understudied. In this work, we leverage long-term operational data, collected from a national-scale production HPC datacenter, and statistically compare how ML and generic jobs can impact the performance, failures, resource utilization, and energy consumption of HPC datacenters. Our study provides key insights, e.g., ML-related power usage causes GPU nodes to run into temperature limitations, median/mean runtime and failure rates are higher for ML jobs than for generic jobs, both ML and generic jobs exhibit highly variable arrival processes and resource demands, significant amounts of energy are spent on unsuccessfully terminating jobs, and concurrent jobs tend to terminate in the same state. We open-source our cleaned-up data traces on Zenodo (https://doi. org/10.5281/zenodo.13685426), and provide our analysis toolkit as software hosted on GitHub (https://github.com/atlarge-research/2024-icpads-hpc-workload-characterization). This study offers multiple benefits for data center administrators, who can improve operational efficiency, and for researchers, who can further improve system designs, scheduling techniques, etc.

crossanalysis↗

ATLAS-MAP: An Automated Test Station for Gated Electronic Transport Measurements

The diversification of electronic materials in devices provides a strong incentive for methods to rapidly correlate device performance with fabrication decisions. In this work, we present a low-cost automated test station for gated electronic transport measurements of field-effect transistors. Utilizing open-source PyMeasure libraries for transparent instrument control, the “ATLAS-MAP” system serves as a customizable interface between sourcemeters and samples under test and is programmed to conduct transfer curve and van der Pauw methods with static and sweeping gate voltages. Zinc oxide transistors of variable thickness (5, 10, and 20 nm) and channel size (50 μm to 3 mm, of equal length and width) were fabricated to validate the design. Standardization of testing procedures and raw data formatting enabled automated data analysis. A detailed list of parts and code files for the system are provided.

36 MATERIALS SCIENCE↗

Physical discovery in representation learning via conditioning on prior knowledge

Recent advances in electron, scanning probe, optical, and chemical imaging and spectroscopy yield bespoke data sets containing the information of structure and functionality of complex systems. In many cases, the resulting data sets are underpinned by low-dimensional simple representations encoding the factors of variability within the data. The representation learning methods seek to discover these factors of variability, ideally further connecting them with relevant physical mechanisms. However, generally, the task of identifying the latent variables corresponding to actual physical mechanisms is extremely complex. Here, we present an empirical study of an approach based on conditioning the data on the known (continuous) physical parameters and systematically compare it with the previously introduced approach based on the invariant variational autoencoders. The conditional variational autoencoder (cVAE) approach does not rely on the existence of the invariant transforms and hence allows for much greater flexibility and applicability. Interestingly, cVAE allows for limited extrapolation outside of the original domain of the conditional variable. However, this extrapolation is limited compared to the cases when true physical mechanisms are known, and the physical factor of variability can be disentangled in full. We further show that introducing the known conditioning results in the simplification of the latent distribution if the conditioning vector is correlated with the factor of variability in the data, thus allowing us to separate relevant physical factors. We initially demonstrate this approach using 1D and 2D examples on a synthetic data set and then extend it to the analysis of experimental data on ferroelectric domain dynamics visualized via piezoresponse force microscopy.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Learning continuous scattering length density profiles from neutron reflectivities using convolutional neural networks

Interpreting neutron reflectivity (NR) data using ad hoc multi-layer models and physics-based models provides information about spatially resolved neutron scattering length density (NSLD) profiles. Recent improvements in data acquisition systems have allowed acquiring thousands of NR curves in a couple of hours, which has led to a need for automated data analysis tools to interpret NR measurements in real-time. Here, we present a machine learning analysis workflow that uses a series of models, based on a convolutional neural network (CNN), to learn the relation between the NSLDs and the NRs, and subsequently produce continuous NSLD profiles directly from NRs. The usefulness of our CNN-based models is demonstrated by constructing NSLDs from NRs of several films containing homopolymer polyzwitterions and diblock copolymers mixed with different types of salts. Comparisons of the NSLDs with those constructed using ad hoc multi-layer models reveal a very good agreement, suggesting the potential of CNN-based models for real-time automated data analysis of NRs.

36 MATERIALS SCIENCE↗

Compactly‐Supported Nonstationary Kernels for Computing Exact Gaussian Processes on Big Data

The Gaussian process (GP) is a widely used method for analyzing large-scale data sets, including spatio-temporal measurements of nonlinear processes that are now commonplace in the environmental sciences. Traditional implementations of GPs involve stationary kernels (also termed covariance functions) that limit their flexibility, and exact methods for inference that prevent application to data sets with more than about 10,000 points. Modern approaches to address stationarity assumptions generally fail to accommodate large data sets, while all attempts to address scalability focus on approximating the Gaussian likelihood, which can involve subjectivity and lead to inaccuracies. In this work, we explicitly derive an alternative kernel that can discover and encode both sparsity and nonstationarity. We embed the kernel within a fully Bayesian GP model and leverage high-performance computing resources to enable the analysis of massive data sets. We demonstrate the favorable performance of our novel kernel relative to existing exact and approximate GP methods across a variety of synthetic data examples. Furthermore, we conduct space–time prediction based on more than 1 million measurements of daily maximum temperature and verify that our results outperform state-of-the-art methods in the Earth sciences. More broadly, having access to exact GPs that use ultra-scalable, sparsity-discovering, nonstationary kernels allows GP methods to truly compete with a wide variety of machine learning methods.

Gaussian processes↗

Scalable Data Center Capacity for DOE's AI Prototype: A Rapidly Available Gigawatt Data Center for DOE

The multilaboratory Gigawatt Data Center working group was commissioned to identify approaches to rapidly establish federal data centers with scalable capacities up to 1,000 MW. These state-of-the-art facilities will serve as hubs for interdisciplinary collaboration, industry partnerships, and transformative applications of artificial intelligence. The proposed strategic shift includes facilitating multilaboratory collaboration, prioritizing operational efficiency, expanding public–private partnerships, optimizing investments, ensuring long-term contractual flexibility, supporting open science and secure data enclaves, and exploiting high-speed national networks. Owing to their extensive experience and best practices, the US Department of Energy national laboratories are uniquely positioned to lead this initiative. We recommend conducting a feasibility analysis to rapidly identify the optimal sites for this initiative, and the effort will likely involve private industry for design, construction, financing, and operational integration. We also propose establishing multiple geographically diverse sites to ensure energy resilience, high operational reliability, and a diverse user base, thereby effectively addressing the nation’s critical needs.

42 ENGINEERING↗

The microbiologist's guide to metaproteomics

Metaproteomics is an emerging approach for studying microbiomes, offering the ability to characterize proteins that underpin microbial functionality within diverse ecosystems. As the primary catalytic and structural components of microbiomes, proteins provide unique insights into the active processes and ecological roles of microbial communities. By integrating metaproteomics with other omics disciplines, researchers can gain a comprehensive understanding of microbial ecology, interactions, and functional dynamics. This review, developed by the Metaproteomics Initiative (www.metaproteomics.org), serves as a practical guide for both microbiome and proteomics researchers, presenting key principles, state-of-the-art methodologies, and analytical workflows essential to metaproteomics. Topics covered include experimental design, sample preparation, mass spectrometry techniques, data analysis strategies, and statistical approaches.

bioinformatics↗

Are light curve classification metrics good proxies for SN Ia cosmological constraining power?

Context. When selecting a light curve classifier for use as part of a photometric supernova Ia (SN Ia) cosmological analysis, it is common to make decisions based on metrics of classification performance, such as the contamination within the photometrically classified SN Ia sample, rather than a measure of cosmological constraining power. If the former is an appropriate proxy for the latter, this practice would eliminate the computational expense of a full cosmology forecast in the analysis pipeline design process. Aims. This study tests the assumption that light curve classification metrics are an appropriate proxy for cosmology metrics. Methods. We emulated photometric SN Ia cosmology light curve samples with controlled contamination rates of individual contaminant classes and evaluated each of them under a set of classification metrics. We then derived cosmological parameter constraints from all samples under two common analysis approaches and quantified the impact of contamination by each contaminant class on the resulting cosmological parameter estimates. Results. We observe that cosmology metrics are sensitive to both the contamination rate and the class of the contaminating population, whereas the classification metrics are shown to be insensitive to the latter. Conclusions. Based on these findings, we discourage any exclusive reliance on light curve classification-based metrics for analysis design decisions, which (counterintuitively) include but are not limited to the classifier choice. Instead, we recommend optimising science analysis pipeline design choices using a metric of the information gained about the physical parameters of interest.

79 ASTRONOMY AND ASTROPHYSICS↗

ESnet Data and AI Workshop Report

In February 2025, the DOE user facility Energy Sciences Network (ESnet) held a three-day Data and AI Workshop in Berkeley, California. The objective of the workshop was to identify challenges within ESnet that could be addressed through data-driven methods, to help define ESnet’s data-analysis requirements, and to shape its AI strategy, guiding data-stewardship efforts and the direction of AI research and AIOps exploration for ESnet7, the next iteration of ESnet’s network. This report summarizes the multi-faceted discussions and findings and presents a set of recommendations for next steps.

97 MATHEMATICS AND COMPUTING↗

Tractometry of the Human Connectome Project: resources and insights

The Human Connectome Project (HCP) has become a keystone dataset in human neuroscience, with a plethora of important applications in advancing brain imaging methods and an understanding of the human brain. We focused on tractometry of HCP diffusion-weighted MRI (dMRI) data. We used an open-source software library (pyAFQ; https://yeatmanlab.github.io/pyAFQ) to perform probabilistic tractography and delineate the major white matter pathways in the HCP subjects that have a complete dMRI acquisition (n = 1,041). We used diffusion kurtosis imaging (DKI) to model white matter microstructure in each voxel of the white matter, and extracted tract profiles of DKI-derived tissue properties along the length of the tracts. We explored the empirical properties of the data: first, we assessed the heritability of DKI tissue properties using the known genetic linkage of the large number of twin pairs sampled in HCP. Second, we tested the ability of tractometry to serve as the basis for predictive models of individual characteristics (e.g., age, crystallized/fluid intelligence, reading ability, etc.), compared to local connectome features. To facilitate the exploration of the dataset we created a new web-based visualization tool and use this tool to visualize the data in the HCP tractometry dataset. Finally, we used the HCP dataset as a test-bed for a new technological innovation: the TRX file-format for representation of dMRI-based streamlines. We released the processing outputs and tract profiles as a publicly available data resource through the AWS Open Data program's Open Neurodata repository. We found heritability as high as 0.9 for DKI-based metrics in some brain pathways. We also found that tractometry extracts as much useful information about individual differences as the local connectome method. We released a new web-based visualization tool for tractometry—“Tractoscope” (https://nrdg.github.io/tractoscope). We found that the TRX files require considerably less disk space-a crucial attribute for large datasets like HCP. In addition, TRX incorporates a specification for grouping streamlines, further simplifying tractometry analysis.

59 BASIC BIOLOGICAL SCIENCES↗

Validation of a Global Geospace Model With a Systems Science Approach Based on Canonical Correlation Analysis

A systems science approach based on canonical correlation analysis (CCA) is applied as a new, behavioral way to validate global geospace models. The biggest novelty of the technique is that it validates models at a system level, whereby a side‐by‐side comparison is performed of CCA applied to a 30‐day observational and the corresponding simulation data sets comprising quiet, moderate and active times. The simulation used the Multiscale Atmosphere‐Geospace Environment (MAGE) model. It is shown that (a) CCA must be combined with sensitivity analysis to be effective, (b) the MAGE model generally reproduces the observed behavior (more so for quieter time intervals), quantified by the intercorrelations between different variables and (c) the technique identifies the SuperMAG SML index as a quantity for which refinements of the model are needed.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Opportunities for Earth Observation to Inform Risk Management for Ocean Tipping Points

Abstract As climate change continues, the likelihood of passing critical thresholds or tipping points increases. Hence, there is a need to advance the science for detecting such thresholds. In this paper, we assess the needs and opportunities for Earth Observation (EO, here understood to refer to satellite observations) to inform society in responding to the risks associated with ten potential large-scale ocean tipping elements: Atlantic Meridional Overturning Circulation; Atlantic Subpolar Gyre; Beaufort Gyre; Arctic halocline; Kuroshio Large Meander; deoxygenation; phytoplankton; zooplankton; higher level ecosystems (including fisheries); and marine biodiversity. We review current scientific understanding and identify specific EO and related modelling needs for each of these tipping elements. We draw out some generic points that apply across several of the elements. These common points include the importance of maintaining long-term, consistent time series; the need to combine EO data consistently with in situ data types (including subsurface), for example through data assimilation; and the need to reduce or work with current mismatches in resolution (in both directions) between climate models and EO datasets. Our analysis shows that developing EO, modelling and prediction systems together, with understanding of the strengths and limitations of each, provides many promising paths towards monitoring and early warning systems for tipping, and towards the development of the next generation of climate models.

Wood, Richard A. (ORCID:0000000239609513)↗

Analog Computing for Science

Conventional digital computing faces fundamental physical limits: large scale computing systems already con sume tens of Megawatts of power, Dennard scaling has ended, and data movement costs dominate application performance. Next generation experimental facilities generate data at rates that overwhelm conventional pro cessing and demand real-time analysis at the source. Analog computing, which exploits the continuous dynamics of physical systems to perform computation, promises a transformative path toward orders-of-magnitude gains in energy efficiency and time-to-solution for scientific workloads.

97 MATHEMATICS AND COMPUTING↗

OrthoPhyl—streamlining large-scale, orthology-based phylogenomic studies of bacteria at broad evolutionary scales

Abstract There are a staggering number of publicly available bacterial genome sequences (at writing, 2.0 million assemblies in NCBI's GenBank alone), and the deposition rate continues to increase. This wealth of data begs for phylogenetic analyses to place these sequences within an evolutionary context. A phylogenetic placement not only aids in taxonomic classification but informs the evolution of novel phenotypes, targets of selection, and horizontal gene transfer. Building trees from multi-gene codon alignments is a laborious task that requires bioinformatic expertise, rigorous curation of orthologs, and heavy computation. Compounding the problem is the lack of tools that can streamline these processes for building trees from large-scale genomic data. Here we present OrthoPhyl, which takes bacterial genome assemblies and reconstructs trees from whole genome codon alignments. The analysis pipeline can analyze an arbitrarily large number of input genomes (>1200 tested here) by identifying a diversity-spanning subset of assemblies and using these genomes to build gene models to infer orthologs in the full dataset. To illustrate the versatility of OrthoPhyl, we show three use cases: E. coli/Shigella, Brucella/Ochrobactrum and the order Rickettsiales. We compare trees generated with OrthoPhyl to trees generated with kSNP3 and GToTree along with published trees using alternative methods. We show that OrthoPhyl trees are consistent with other methods while incorporating more data, allowing for greater numbers of input genomes, and more flexibility of analysis.

59 BASIC BIOLOGICAL SCIENCES↗

Galaxy-multiplet clustering from DESI DR2

We present an efficient estimator for higher-order galaxy clustering using small groups of nearby galaxies, or multiplets. Using the Luminous Red Galaxy (LRG) sample from the Dark Energy Spectroscopic Instrument (DESI) Data Release 2, we identify galaxy multiplets as discrete objects and measure their cross-correlations with the general galaxy field. Our results show that the multiplets exhibit stronger clustering bias as they trace more massive dark matter halos than individual galaxies. When comparing the observed clustering statistics with the mock catalogs generated from the N-body simulation AbacusSummit, we find that the mocks underpredict multiplet clustering despite reproducing the galaxy two-point auto-correlation reasonably well. This discrepancy indicates that the standard Halo Occupation Distribution (HOD) model is insufficient to describe the properties of galaxy multiplets, revealing the greater constraining power of this higher-order statistic on galaxy-halo connection and the possibility that multiplets are specific to additional assembly bias. We demonstrate that incorporating secondary biases into the HOD model improves agreement with the observed multiplet statistics, specifically by allowing galaxies to preferentially occupy halos in denser environments. Our results highlight the potential of utilizing multiplet clustering, beyond traditional two-point correlation measurements, to break degeneracies in models describing the galaxy-dark matter connection.

cosmology↗