Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data integrity”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Tethered Balloon System Merged Data (TBSMERGED) Value-Added Product Report

The U.S. Department of Energy Atmospheric Radiation Measurement (ARM) user facility’s Tethered Balloon System Merged Data (TBSMERGED) Value-Added Product (VAP) integrates data from instruments flown on ARM’s tethered balloon system missions that collect in situ measurements of temperature, humidity, wind speed, wind direction, and aerosol properties with estimates of cloud base and boundary-layer height from a surface-based ceilometer to improve the ease of use of tethered balloon system (TBS) data sets. The TBSMERGEDINCLOUD VAP, a variation of TBSMERGED, includes supercooled liquid water content (tbsslwc) measurements collected within the cloud.

54 ENVIRONMENTAL SCIENCES↗

DeepLynx Ecosystem 2025

Poor data integration and governance continue to plague complex engineering projects, resulting in missed cost, schedule, and performance targets. Departments operate in isolated systems with manual data exchange, creating fragmented information that compounds errors and leads to significant delays and cost overruns. The DeepLynx ecosystem addresses these challenges through an open-source, modular data management platform that transforms fragmented project data into an integrated digital thread. Built on a federated microservice architecture, the ecosystem comprises seven specialized tools centered around DeepLynx Nexus, a unified data catalog with hierarchical organization and graph-based navigation capabilities. The ecosystem includes: DeepLynx Stream for real-time timeseries data ingestion from industrial sources; DeepLynx Ingest for governed data uploads with formal review workflows; DeepLynx Lattice for ontology-based entity and relationship extraction; DeepLynx Run for workflow orchestration and secure AI/ML compute; DeepLynx Visualize for 3D digital twin visualization; and DeepLynx Insight for AI-assisted document analysis with traceable, grounded responses. Deployable in cloud, on-premise, or hybrid environments using containerized Docker applications and Helm charts, the DeepLynx ecosystem provides flexible infrastructure that adapts to organizational requirements. By consolidating project data into a unified data lake with role-based access controls and OAuth2 authentication, DeepLynx enables digital thread and digital twin capabilities that improve decision-making, reduce risk, and support complex engineering workflows throughout the project lifecycle.

42 - ENGINEERING↗

Energy Community Atlas

The Energy Community Atlas provides efficient access to authoritative, curated, and relevant data that is vital to supporting energy planning, development, and economic growth across the U.S. In this effort, researchers at the National Energy Technology Laboratory (NETL) are utilizing advanced data visualization and transformation capabilities to develop an integrated, data atlas and resource focused on supporting energy community transitions to new manufacturing opportunities. Specifically, this project is working to find, acquire, integrate, and virtually host in a user-friendly, public and private solution from available resources, relevant to understanding and characterizing fossil energy communities themselves and inform energy planning, development, and economic growth opportunities, including opportunities for co-development to support manufacturing, critical materials, and more. This Atlas when complete is to offer a one-stop-shop for stakeholders to derive new insights to accelerate energy investments and strategic decision support needs. These are following datasets that are available as part of this ongoing project • Energy Community Atlas Map Package - This is ArcPro Map package and it contains all of the symbolized layers along with ArcPro map and geodatabase • Energy Community Atlas ArcGIS REST service - https://www.arcgis.com/apps/mapviewer/index.html?panel=gallery&suggestField=true&layers=537ced69bd88440380a62c2ec8aca30c • README Energy Community Atlas - Read me word document that has details about feature classes in Map package, ArcPro map and ArcGIS Rest Service

Bipartisan Infrastructure Law↗

Navigating Integration: Key Challenges for Data Centers, Nuclear Stakeholders, and Utility Operators

The rapid expansion of data centers, driven by the exponential growth in data-processing and storage needs, presents significant challenges and opportunities for various stakeholders, including data center developers, nuclear energy providers, and utility companies. Data centers are projected to consume 6.7–12% of United States (U.S.) electricity by 2028, driven by artificial intelligence (AI) and cloud-computing demands. Nuclear energy offers reliability and dispatchable baseload power, but data centers need power now while nuclear still needs time to address siting, fast power ramping, and regulatory hurdles. Utilities must keep pace with the unprecedented acceleration of large load interconnection requests and urgently adapt to high-density loads while maintaining grid stability, reliability, and accelerating interconnection timelines. This report dives into these challenges and proposes key collaboration strategies to streamline data center integration that aligns with recent federal initiatives like America’s AI Action Plan and related executive orders that emphasize the importance of data center growth, nuclear energy expansion, and maintaining a competitive edge in the global AI race.

21 - SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLAN↗

Mass Spectrometry Sample Submission Portal

Each step in the scientific process generates contextual information about the data that is important to consider when performing data integration, developing models of biological process, or training AI models. We will develop a flexible, template-driven tool that will log biological samples, capture metadata about those samples, and track the type(s) of analysis being performed by researchers providing samples for analysis by mass spectrometry.

97 MATHEMATICS AND COMPUTING↗

Machine Learning to Select Experiments Driven by Fundamental Science and Applications for Targeted Nuclear Data Improvement

This work describes a blueprint for a process that accelerates progress in science by quantitatively answering the following question: What is the optimal combination of fundamental-science and application-driven experiments to maximally reduce pertinent data uncertainties? Answering this question entails solving a high-dimensional and complex optimization problem that is best solved with advanced statistic techniques often classified as machine learning. We apply this process within the framework of nuclear data with the aim to select an experiment combination that will reduce uncertainties in 239 Pu nuclear data for neutron energies between 1 and 600 keV. In this field, fundamental-physics driven data, called differential, look at one nuclear physics observable at a time. They are contrasted to application-driven, integral, data where one or few resulting values inform a broad set of nuclear data across several nuclides and energies. The candidates for integral experiments are criticality measurements that were refined by a genetic algorithm to be maximally sensitive to 239 Pu fission cross sections in the desired energy range. Twenty-three candidate differential experiments were investigated and span multiple nuclear physics observables (e.g., total, capture cross sections) for isotopes appearing in the integral experiments. The optimal combination among these candidate experiments was investigated via generalized least squares fitting, augmented with Gaussian processes to ameliorate statistical irregularities in data, and the D-optimality criterion. The latter evaluates for each pair of candidates the joint reduction in uncertainties of all 12200 nuclear data appearing in the integral experiments compared to the knowledge we have from 168 past experiments, theory, and nuclear data. We chose as differential measurements those that investigate 63 Cu and 239 Pu total cross sections, based on D-optimality rank and feasibility constraints. Two integral (criticality) experiments were selected: An experiment with Al 2 ⁢O 3 and graphite interleaved with Pu and a thick Cu reflector explores 1–30 keV, while we target the 30–600 keV range with an experiment that swaps boron in place of graphite with a different geometry.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Investigation of Benchmark $k$ eff Sensitivity and Uncertainty for 239 Pu fission in Specific Energy Ranges

Nuclear data at intermediate energies (from 1 to 100s of keV) are evaluated based on scarce differential data and theory unable to capture physics’ expected structure. There is also a lack of integral data. This is a known deficiency and is challenging to address. Calculated effective multiplication factor, k eff , values for intermediate energy experiments are ~25× further from experiment than for fast energies and are often well outside the experimental uncertainties. The goal of the PARADIGM (PARallel Approach of Differential and InteGral Measurements) project is to significantly re duce the uncertainties of intermediate energy nuclear data for 239 Pu. To this end, PARADIGM simultaneously optimizes experiments at both the Los Alamos Neutron Science Center (LANSCE) and National Criticality Experiments Research Center (NCERC). The combined set of data will inform new intermediate-energy nuclear data. By execution of differential and integral experiments, establishment of new theory, and undertaking nuclear data evaluation in parallel, the timeline to deliver improved nuclear data to users will be reduced significantly that is to three years. For the PARADIGM project, it was decided to optimize an integral experiment for two neutron energy ranges, within the full intermediate energy range. The low energy range goes from 1 to 30 keV, while the higher energy range goes from 30 to 600 keV. This work focuses on nuclear data sensitivities and uncertainties for 239 Pu fission for existing experiments in the International Criticality Safety Benchmark Evaluation Project (ICSBEP). When designing new experiments, it is important to understand what benchmarks currently exist. For a more traditional experiment design (in which a specific application model(s) exists), comparisons would be made between the application model(s) and existing benchmarks. For PARADIGM, there is no specific application model, but instead the specific nuclear data reaction and energy ranges of interest can be explored for existing benchmarks.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Tools for Visualization and Analysis of Small-Angle Neutron Scattering Data: Descriptions and Examples

A great deal of progress has been made in improving the data reduction experience for the SANS instruments at the SNS and HFIR at ORNL. The existing data reduction toolset, drtsans, makes it possible to integrate data analysis and visualization tools into the data reduction scripts, thereby providing new opportunities for more automated data processing for users of the SNS and HFIR. Here, the first set of tools developed is described with usage examples.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Integrating Multi-Source Data for Bi-Level Traffic Simulator Calibration: A Literature Review and Highway Case Study

Traffic simulation serves as a powerful tool for pre-evaluating policies and technologies. In this context, simulation-based Dynamic traffic assignment (DTA) models are capable of capturing traffic dynamics. They are well-known as critical tools in controlling and predicting traffic situations. The reliability of simulation results heavily depends on the calibration process. Most studies in the literature formulate and calibrate simulators based on a single source of collected data or multiple data sets with the same spatiotemporal characteristics. However, in practice, traffic data is collected by various tools with usually different spatial and temporal resolutions. This study introduces a novel approach to taking into account diverse input data from a variety of sources. An iterative bi-level solution is proposed. to equally treat traffic flow and speed data. The upper level solves flow calibration with the exact solution method, and the lower level calibrates the speed with the simultaneous perturbation stochastic approximation (SPSA) algorithm. Subsequently, the effectiveness of the proposed model is investigated using data from a six-mile section of Nashville's I-24 highway in Tennessee. The results demonstrate that our proposed model creates an effective feedback loop between the optimizer and the simulator for calibrating flow and speed to reduce the error between simulated and real data.

42 ENGINEERING↗

Online and Offline Data Quality Monitoring for the Mu2e Calorimeter

This thesis presents the design, implementation, and validation of a calorimeter Data Quality Monitoring (DQM) toolchain for the Mu2e experiment at Fermilab. Mu2e searches for charged lepton flavor violation via coherent muon-to-electron conversion in the field of an aluminum nucleus, $\mu^- Al \rightarrow e^-Al$, a process whose observation would constitute clear evidence of physics beyond the Standard Model. Achieving target sensitivity requires stringent control of detector performance and data integrity during acquisition, as subtle issues in readout configuration, data formatting, or electronics behavior can compromise reconstruction and bias downstream analyzes. To address these challenges, this work develops a multi-layer DQM approach spanning both raw data validation and reconstructed digi-level diagnostics. At the low level, a fragment analysis component performs word- and bit-field decoding of calorimeter readout blocks, enabling sanity checks of the expected structure and producing detailed error and integrity statistics useful for commissioning and troubleshooting. At the digi level, the CaloDigiDQM analyzer is implemented within the art framework and transforms each CaloDigiCollection into a structured hierarchy of ROOT histograms designed for fast drill-down diagnostics. The module generates coherent monitoring views at global, disk, board, and channel granularity, including occupancy, waveform-derived features (baseline, RMS, peak amplitude and position), and left-right sensor consistency metrics. Detector-aware channel-to-electronics mapping is performed through the conditions system (CaloDAQMap), ensuring that diagnostics remain aligned with hardware identifiers used in operations. For end-to-end testing without reliance on live DAQ data, a synthetic CaloDigi producer is developed to generate realistic waveforms with controlled noise and pulse shapes. The resulting system supports both offline ROOT-file production and online operation, including optional histogram streaming through otsdaq via ots::HistoSender. This toolchain provides a practical and scalable foundation for calorimeter commissioning and stable data collection, enabling early detection of anomalies and reducing operational risk for Mu2e.

Vakulenko, Mark [Drew U.] (ORCID:0009000276197818)↗

Automated pipeline processing X-ray diffraction data from dynamic compression experiments on the Extreme Conditions Beamline of PETRA III

Presented and discussed here is the implementation of a software solution that provides prompt X-ray diffraction data analysis during fast dynamic compression experiments conducted within the dynamic diamond anvil cell technique. It includes efficient data collection, streaming of data and metadata to a high-performance cluster (HPC), fast azimuthal data integration on the cluster, and tools for controlling the data processing steps and visualizing the data using the DIOPTAS software package. This data processing pipeline is invaluable for a great number of studies. The potential of the pipeline is illustrated with two examples of data collected on ammonia–water mixtures and multiphase mineral assemblies under high pressure. The pipeline is designed to be generic in nature and could be readily adapted to provide rapid feedback for many other X-ray diffraction techniques, e.g. large-volume press studies, in situ stress/strain studies, phase transformation studies, chemical reactions studied with high-resolution diffraction etc.

97 MATHEMATICS AND COMPUTING↗

Rock Physics-Based Data Assimilation of Integrated Continuous Active-Source Seismic and Pressure Monitoring Data during Geological Carbon Storage

Summary There has been substantial controversy concerning the role of geological carbon storage (GCS) in sequestering anthropogenic carbon emissions to mitigate climate change and global warming. Arguments center on the inability to monitor a geological storage site precisely and continuously, especially highlighting the associated costs and spatiotemporal trade-offs when using conventional subsurface monitoring techniques (well logs, core samples, chemical tracers, and 4D seismics). Active surveillance of GCS sites is essential for managing and mitigating potential leaks but is also required by regulation. With the goal of enhancing the monitoring capability at GCS sites, we present a rock physics-based joint data assimilation model to study a popular GCS site at Cranfield, Mississippi, USA. Synthetic continuous active-source seismic monitoring (CASSM) data (in the form of Vp and Qp measurements) and wellbore pressure monitoring data are assimilated with an ensemble of reservoir realizations to monitor gas saturation and reservoir pressure changes over a period of 100 years. Synthetic seismic attributes are generated using rock physics models (RPMs) and wellbore pressure monitoring data are extracted from the ground truth. Two assimilation methods, ensemble Kalman filter (EnKF) and ensemble Kalman smoother (EnKS), are tested in an observation system simulation experiment (OSSE) environment to assess the prediction accuracy of the individual and composite observation systems. The joint monitoring system achieves more accurate estimates of gas saturation and pressure, across the time span from start of injection to end of forecast, as compared to a single type of monitoring tool and irrespective of data assimilation algorithm choice. These results indicate that jointly assimilated data from two types of sensors (in this case, crosswell seismic and downhole pressure) may lead to a more risk-reducing monitoring design. One would expect that more data, vis-à-vis inclusion of a new sensor type, will improve the accuracy of any GCS monitoring system. However, from a practical standpoint, one important question is whether such a gain in accuracy is worth the additional cost associated with the new sensor. This paper focuses on quantifying the gain in accuracy, such that a practitioner can answer this question.

Engineering↗

A Nuclear Security Enterprise Study of High-Reliability Systems, Collaboration, and Data

It may seem simple and trivial, but defining the difference between data and information is contested and has implications that may affect the security of United States interests and even cost lives. For security, data are raw facts or figures without context, while information is the compilation or articulation of data that forms context. Security depends on clarity in the differences between data and information and controlling them. Control is necessary to ensure that data and information are not inadvertently released to foreign governments, the public, or those without Need-to-Know. A primary concern in the practice of security is the control of data to avoid the inadvertent conversion to sensitive information. The complexity of this concern is further augmented when institutions are part of tightly coupled networks that informally share data and information. Additionally, those that share data as a function of legislative action—and/or formally integrate data and information system infrastructures—may be a higher security risk. This paper will present a case study that utilizes elements of literature from Knowledge Management and networks to tell a story of an issue in security—specifically, controlling the conversion of data to information.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

CrossMP: Enabling Cross-Modality Translation between Single-Cell RNA-Seq and Single-Cell ATAC-Seq through Web-Based Portal

In recent years, there has been a growing interest in profiling multiomic modalities within individual cells simultaneously. One such example is integrating combined single-cell RNA sequencing (scRNA-seq) data and single-cell transposase-accessible chromatin sequencing (scATAC-seq) data. Integrated analysis of diverse modalities has helped researchers make more accurate predictions and gain a more comprehensive understanding than with single-modality analysis. However, generating such multimodal data is technically challenging and expensive, leading to limited availability of single-cell co-assay data. Here, we propose a model for cross-modal prediction between the transcriptome and chromatin profiles in single cells. Our model is based on a deep neural network architecture that learns the latent representations from the source modality and then predicts the target modality. It demonstrates reliable performance in accurately translating between these modalities across multiple paired human scATAC-seq and scRNA-seq datasets. Additionally, we developed CrossMP, a web-based portal allowing researchers to upload their single-cell modality data through an interactive web interface and predict the other type of modality data, using high-performance computing resources plugged at the backend.

59 BASIC BIOLOGICAL SCIENCES↗

Expanding the access of wearable silicone wristbands in community-engaged research through best practices in data analysis and integration

Wearable silicone wristbands are a rapidly growing exposure assessment technology that offer researchers the ability to study previously inaccessible cohorts and have the potential to provide a more comprehensive picture of chemical exposure within diverse communities. However, there are no established best practices for analyzing the data within a study or across multiple studies, thereby limiting impact and access of these data for larger meta-analyses. We utilize data from three studies, from over 600 wristbands worn by participants in New York City and Eugene, Oregon, to present a first-of-its-kind manuscript detailing wristband data properties. We further discuss and provide concrete examples of key areas and considerations in common statistical modeling methods where best practices must be established to enable meta-analyses and integration of data from multiple studies. Finally, we detail important and challenging aspects of machine learning, meta-analysis, and data integration that researchers will face in order to extend beyond the limited scope of individual studies focused on specific populations.

Bramer, Lisa M.↗

Data to Accompany: Expanding the access of wearable silicone wristbands in community-engaged research through best practices in data analysis and integration

Wearable silicone wristbands are a rapidly growing exposure assessment technology that offer researchers the ability to study previously inaccessible cohorts and have the potential to provide a more comprehensive picture of chemical exposure within diverse communities. However, there are no established best practices for analyzing the data within a study or across multiple studies, thereby limiting impact and access of these data for larger meta-analyses. We utilize data from three studies, from over 600 wristbands worn by participants in New York City and Eugene, Oregon, to present a first-of-its-kind manuscript detailing wristband data properties. We further discuss and provide concrete examples of key areas and considerations in common statistical modeling methods where best practices must be established to enable meta-analyses and integration of data from multiple studies. Finally, we detail important and challenging aspects of machine learning, meta-analysis, and data integration that researchers will face in order to extend beyond the limited scope of individual studies focused on specific populations.

Bramer, Lisa M↗

Genome-Resolved Metaproteomics Decodes the Microbial and Viral Contributions to Coupled Carbon and Nitrogen Cycling in River Sediments

Rivers have a significant role in global carbon and nitrogen cycles, serving as a nexus for nutrient transport between terrestrial and marine ecosystems. Although rivers have a small global surface area, they contribute substantially to worldwide greenhouse gas emissions through microbially mediated processes within the river hyporheic zone. Despite this importance, research linking microbial and viral communities to specific biogeochemical reactions is still nascent in these sediment environments. To survey the metabolic potential and gene expression underpinning carbon and nitrogen biogeochemical cycling in river sediments, we collected an integrated data set of 33 metagenomes, metaproteomes, and paired metabolomes. We reconstructed over 500 microbial metagenome-assembled genomes (MAGs), which we dereplicated into 55 unique, nearly complete medium- and high-quality MAGs spanning 12 bacterial and archaeal phyla. We also reconstructed 2,482 viral genomic contigs, which were dereplicated into 111 viral MAGs (vMAGs) of >10 kb in size. As a result of integrating gene expression data with geochemical and metabolite data, we created a conceptual model that uncovered new roles for microorganisms in organic matter decomposition, carbon sequestration, nitrogen mineralization, nitrification, and denitrification. We show how these metabolic pathways, integrated through shared resource pools of ammonium, carbon dioxide, and inorganic nitrogen, could ultimately contribute to carbon dioxide and nitrous oxide fluxes from hyporheic sediments. Further, by linking viral MAGs to these active microbial hosts, we provide some of the first insights into viral modulation of river sediment carbon and nitrogen cycling.

54 ENVIRONMENTAL SCIENCES↗