Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data enrichment”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Data-Driven Optimization of Pixelated CdZnTe Spectrometers for Uranium Enrichment Assay

Here, in recent work [Vavrek et al. (2025)], we developed the performance optimization framework spectre-ml for gamma spectrometers with variable performance across many readout channels. The framework uses non-negative matrix factorization (NMF) and clustering to learn groups of similarly-performing channels and sweep through various learned channel combinations to optimize the performance tradeoff of including worse-performing channels for better total efficiency. In this work, we integrate the pyGEM uranium enrichment assay code with our spectre-ml framework, and show that the U-235 enrichment relative uncertainty can be directly used as an optimization target. We find that this optimization reduces relative uncertainties after a 30 -minute measurement by an average of 20%, as tested on six different H3D M400 CdZnTe spectrometers, which can significantly improve uranium non-destructive assay measurement times in nuclear safeguards contexts. Additionally, this work demonstrates that the spect re-ml optimization framework can accommodate arbitrary end-user spectroscopic analysis code and performance metrics, enabling future optimizations for complex Pu spectra.

Gamma-ray detection↗

Data for Production of Designer Xylose-Acetic Acid Enriched Hydrolysate from Bioenergy Sorghum, Oilcane, and Energycane Bagasses

Xylan accounts for up to 40% of the structural carbohydrates in lignocellulosic feedstocks. Along with xylan, acetic acid in sources of hemicellulose can be recovered and marketed as a commodity chemical. Through vibrant bioprocessing innovations, converting xylose and acetic acid into high-value bioproducts via microbial cultures improves the feasibility of lignocellulosic biorefineries. Enzymatic hydrolysis using xylanase supplemented with acetylxylan esterase (AXE) was applied to prepare xylose-acetic acid enriched hydrolysates from bioenergy sorghum, oilcane, or energycane using sequential hydrothermal-mechanical pretreatment. Various biomass solids contents (15 to 25%, w/v) and xylanase loadings (140 to 280 FXU/g biomass) were tested to maximize xylose and acetic acid titers. The xylose and acetic acid yields were significantly improved by supplementing with AXE. The optimal yields of xylose and acetic acid were 92.29% and 62.26% obtained from hydrolyzing energycane and oilcane at 25% and 15% w/v biomass solids using 280 FXU xylanase/g biomass and AXE, respectively.

Biomass Analytics↗

As-Built Simulation of the High Flux Isotope Reactor

The Oak Ridge National Laboratory High Flux Isotope Reactor (HFIR) is an 85 MWt flux trap-type research reactor that supports key research missions, including isotope production, materials irradiation, and neutron scattering. The core consists of an inner and an outer fuel element containing 171 and 369 involute-shaped plates, respectively. The thin fuel plates consist of a U 3 O 8 -Al dispersion fuel (highly enriched), an aluminum-based filler, and aluminum cladding. The fuel meat thickness is varied across the width of the involute plate to reduce thermal flux peaks at the radial edges of the fuel elements. Some deviation from the designed fuel meat shaping is allowed during manufacturing. A homogeneity scan of each fuel plate checks for potential anomalies in the fuel distribution by scanning the surface of the plate and comparing the attenuation of the beam to calibration standards. While typical HFIR simulations use homogenized fuel regions, explicit models of the plates were developed under the Low-Enriched Uranium Conversion Program. These explicit models typically include one inner and one outer fuel plate with nominal fuel distributions, and then the plates are duplicated to fill the space of the corresponding fuel element. Therefore, data extracted from these simulations are limited to azimuthally averaged quantities. To determine the reactivity and physics impacts of an as-built outer fuel element and generate azimuthally dependent data in the element, 369 unique fuel plate models were generated and positioned. This model generates the three-dimensional (i.e., radial–axial–azimuthal) plate power profile, where the azimuthal profile is impacted by features within the adjacent control element region and beryllium reflector. For an as-built model of the outer fuel element, plate-specific homogeneity data, 235 U loading, enrichment, and channel thickness measurements were translated into the model, yielding a much more varied azimuthal power profile encompassed by uncertainty factors in analyses. These models were run with the ORNL-TN and Shift Monte Carlo tools, and they contained upwards of 500,000 cells and 100,000 unique tallies.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

KBase Narrative - Supplemental Data for "Mixed waste contamination selects for a mobile genetic element population enriched in multiple heavy metal resistance genes"

Here we provide the complete set of circular elements produced by SCAPP from metagenomic data that may represent mobile genetic elements (MGEs). For the publication, these sequences were de-replicated within the high [U] and low [U] sample sets. Additionally, 98 of the circular elements were removed for downstream analysis due to high suspicion of being (1) chloroplast or mitochondrial sequences or (2) erroneously circularized long repeat regions. The complete list of circular elements used for the analysis along with metadata can be found in Tables S5 and S6.

Goff, Jennifer↗

A semidominant point mutation of Mediator tail subunit MED5b in Arabidopsis leads to altered enrichment of H3K27me3 and reduced expression of targets of MYC2

Abstract The Mediator complex coordinates regulatory input for transcription driven by RNA polymerase II in eukaryotes. reduced epidermal fluorescence4-3 (ref4-3) is a semidominant mutation that results in a single amino acid substitution in the Mediator tail subunit Med5b. Previous characterization of ref4-3 revealed altered expression of a variety of loci in Arabidopsis, including those contributing to phenylpropanoid biosynthesis. Examination of existing RNA-seq data indicated that loci enriched for the transcriptionally repressive chromatin modification H3K27me3 are overrepresented among genes that are misregulated in ref4-3. We used ChIP-seq and RNA-seq to examine the possibility that perturbation of H3K27me3 homeostasis in ref4-3 plants contributed to altered transcript levels. We observed that ref4-3 results in a modest global reduction of H3K27me3 at enriched loci and that this reduction is not dependent on gene expression; however, altered H3K27me3 was not strongly predictive of altered expression in ref4-3 plants. Instead, our analyses revealed a substantial enrichment of targets of the MYC2 transcriptional regulator among genes that exhibit decreased expression in ref4-3. Consistent with previous characterization of ref4-3, we observed that ref4-3-dependent decreased expression of MYC2 targets can be suppressed by loss of another Mediator tail subunit, MED25. This observation is consistent with previous biochemical characterization of MYC2. Our data highlight the diverse and distinct impacts that a single amino acid change in the tail subunit of Mediator can have on transcriptional circuits and raise the prospect that Mediator directly contributes to H3K27me3 homeostasis in plants.

Long, Jiaxin (ORCID:0000000335150706)↗

Proteogenomic characterization of difficult-to-treat breast cancer with tumor cells enriched through laser microdissection

Abstract Background Breast cancer (BC) is the most commonly diagnosed cancer and the leading cause of cancer death among women globally. Despite advances, there is considerable variation in clinical outcomes for patients with non-luminal A tumors, classified as difficult-to-treat breast cancers (DTBC). This study aims to delineate the proteogenomic landscape of DTBC tumors compared to luminal A (LumA) tumors. Methods We retrospectively collected a total of 117 untreated primary breast tumor specimens, focusing on DTBC subtypes. Breast tumors were processed by laser microdissection (LMD) to enrich tumor cells. DNA, RNA, and protein were simultaneously extracted from each tumor preparation, followed by whole genome sequencing, paired-end RNA sequencing, global proteomics and phosphoproteomics. Differential feature analysis, pathway analysis and survival analysis were performed to better understand DTBC and investigate biomarkers. Results We observed distinct variations in gene mutations, structural variations, and chromosomal alterations between DTBC and LumA breast tumors. DTBC tumors predominantly had more mutations inTP53,PLXNB3, Zinc finger genes, and fewer mutations inSDC2,CDH1,PIK3CA,SVIL, andPTEN. Notably, Cytoband 1q21, which contains numerous cell proliferation-related genes, was significantly amplified in the DTBC tumors. LMD successfully minimized stromal components and increased RNA–protein concordance, as evidenced by stromal score comparisons and proteomic analysis. Distinct DTBC and LumA-enriched clusters were observed by proteomic and phosphoproteomic clustering analysis, some with survival differences. Phosphoproteomics identified two distinct phosphoproteomic profiles for high relapse-risk and low relapse-risk basal-like tumors, involving several genes known to be associated with breast cancer oncogenesis and progression, includingKIAA1522,DCK,FOXO3,MYO9B,ARID1A,EPRS,ZC3HAV1, andRBM14. Lastly, an integrated pathway analysis of multi-omics data highlighted a robust enrichment of proliferation pathways in DTBC tumors. Conclusions This study provides an integrated proteogenomic characterization of DTBC vs LumA with tumor cells enriched through laser microdissection. We identified many common features of DTBC tumors and the phosphopeptides that could serve as potential biomarkers for high/low relapse-risk basal-like BC and possibly guide treatment selections.

Oncology↗

Integral Experiment Request 523 CED-2 Report

This report documents the final design phase of the Critical Experiment Design (CED-2) conducted as part of integral experiment request (IER) 523. The purpose of IER 523 is to determine critical configurations of 35 weight percent (wt%) enriched uranium dioxide beryllium oxide (UO 2 -BeO) material driven by an annular ring of Seven Percent Critical Experiment (7uPCX) fuel rods at Sandia National Laboratories (Sandia). The experiments will provide benchmark data on water moderated, intermediately enriched UO2 systems as well as Be nuclear data. The experiment will also provide partial validation for the beryllium oxide (BeO) thermal neutron scattering law (TSL) in the thermal energy range. Experiment design concepts, neutronic analysis results, and proposed paths for continuing the CED process are presented. This report builds on the feasibility and justification of experimental need report (CED-0) and preliminary experiment design report (CED-1) completed in December 2021 and September 2023, respectively [1, 2].

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Nuclear Data–Induced Uncertainties in Criticality Safety Analyses for High-Burnup and Extended Enrichment Fuels

Criticality safety analyses are conducted to show compliance with regulatory standards and to demonstrate safe operational conditions during the storage and transportation of spent nuclear fuel. Given the increased interest in the industry in low-enriched uranium plus (LEU+) and higher-burnup fuel, it is important to study the impact of such fuels’ use on criticality safety analyses and the resulting nuclear data–induced uncertainties. Here, in this work, nominal pressurized water reactor assemblies with LEU+ fuel enrichments up to 8 wt% 235 U and high burnups up to 80 GWd/tonne U were studied. The assemblies were placed in a generic burnup credit cask GBC-32. As a result of the different covariance libraries, using the ENDF/B-VII.1 nuclear data library consistently resulted in lower nuclear data uncertainties than did the use of the ENDF/B-VIII.0 data library. The highest contribution in the nuclear data–induced uncertainties resulted from the major actinides, and their contribution increased with increasing burnup and enrichment.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Solving differential equations using deep neural networks

Recent work on solving partial differential equations (PDEs) with deep neural networks (DNNs) is presented. The paper reviews and extends some of these methods while carefully analyzing a fundamental feature in numerical PDEs and nonlinear analysis: irregular solutions. First, the Sod shock tube solution to the compressible Euler equations is discussed and analyzed. This analysis includes a comparison of a DNN-based approach with conventional finite element and finite volume methods, and demonstrates that the DNN is competitive in terms of degrees of freedom required for a given accuracy. Further, the DNN-based approach is extended to consider performance improvements and simultaneous parameter space exploration. Next, a shock solution to compressible magnetohydrodynamics (MHD) is solved for, and used in a scenario where experimental data is utilized to enhance a PDE system that is a priori insufficient to validate against the observed/experimental data. This is accomplished by enriching the model PDE system with source terms that are then inferred via supervised training with synthetic experimental data. The resulting DNN framework for PDEs enables straightforward system prototyping and natural integration of large data sets (be they synthetic or experimental), all while simultaneously enabling single-pass exploration of an entire parameter space.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Deciphering the Distribution and Crystal-Chemical Environment of Arsenic, Lead, Silica, Phosphorus, Tin, and Zinc in a Porous Ferrihydrite Grain Using Transmission Electron Microscopy and Atom Probe Tomography

Here the interaction of contaminants and nutrients with soil constituents is controlled by processes in intergranular and intragranular pore spaces of organic matter or/and common secondary minerals such as ferrihydrite, ~Fe 3+ 10 O 14 (OH) 2 . This contribution shows that distribution and clustering of the contaminants As, P, Pb, Si, Sn, and Zn in a porous ferrihydrite grain is greatly affected by the heterogeneous size distribution and chemical composition of the pores as well as the ability of their polyhedra to polymerize with the same type of polyhedron. Transmission electron microscopy (TEM) and atom probe tomography (APT) studies are conducted on focused ion beam (FIB) sections extracted from a porous ferrihydrite grain from the smelter-impacted topsoil in Sudbury, Ontario, Canada. The ferrihydrite grain has pore spaces ranging in diameter from tens to hundreds of nanometers. TEM and scanning-TEM studies indicate that the surfaces of the pore walls are enriched in Si. APT data in conjunction with First Near Neighbor (1NN) analyses indicate different degrees of clustering of Pb, As, Sn, Zn, Si, and P within the sample and selected domains. Careful evaluations of 3D atomic plots and 1NN distances indicates the occurrence of polymerized arsenite-, silica-, Sn-, and Zn-polyhedra within pore spaces of the ferrihydrite. Deciphering adsorption, polymerization, and nucleation processes in porous Fe-(hydr)oxides and other soil constituents requires multianalytical approaches and, in this regard, we discuss the advantages and disadvantages of the combination of TEM and APT for characterizing complex environmental samples at the atomic to nanometer scale.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

ATIS

Any Threat Intelligence to STIX (ATIS) autogenerates and enriches STIX bundles with data from open source threat intelligence sources.

McCampbell, Taylor↗

Statistical and Machine Learning Approaches to Analyzing Pipeline Incidents in the United States (2010–2024)

This study applies machine learning methods to analyze natural gas pipeline incidents in the United States using the Pipeline and Hazardous Materials Safety Administration (PHMSA) Gas Distribution Incident Dataset (2010–2024). The dataset includes over 600 variables describing incident characteristics, infrastructure attributes, and contributing factors associated with unintentional gas releases. The objective is to assess whether these features can reliably predict the underlying cause of pipeline failures. Multinomial logistic regression and Random Forest models were developed to classify incident causes, including excavation damage, corrosion, equipment failure, and natural forces. Results show that excavation damage is both the most frequent and most predictable cause, with models achieving strong performance for this category. However, when excavation damage is excluded, model accuracy declines significantly, with some models performing near random levels. Across all approaches, severe class imbalance and limited variability in key predictors constrain predictive performance. Pipeline age and diameter emerge as the most influential variables, but they provide insufficient discriminatory power to distinguish among less frequent failure types. These findings indicate that non-excavation-related incidents are rare, heterogeneous, and weakly represented in the dataset, limiting the effectiveness of machine learning classification. Overall, this study highlights the structural limitations of the PHMSA dataset for predictive modeling and underscores the need for improved data balance and feature enrichment. The results reinforce excavation damage prevention as the most impactful strategy for reducing pipeline incidents.

03 NATURAL GAS↗

Corrosion Compatibility of Stainless Steels and Nickel in Pyrolysis Biomass-Derived Oil at Elevated Storage Temperatures

Corrosion compatibility of stainless steels and nickel (Ni200) was assessed in fast pyrolysis bio-oil produced from pyrolysis of high ash and high moisture forest residue biomass. Sample mass change, ICP-MS and post-exposure electron microscopy characterization was used to investigate the extent of corrosion. Among the tested samples, type 430F and type 316 stainless steels (SS430F and SS316) and Ni200 (~98.5% Ni) showed minimal mass changes (less than 2 mg∙cm−2) after the bio-oil exposures at 50 and 80 °C for up to 168 h. SS304 was also considered to be compatible in the bio-oil due to its relatively low mass change (1.6 mg∙cm−2 or lower). SS410 samples showed greater mass loss values even after exposures at a relatively low temperature of 35 °C. Fe/Cr values from ICP-MS data implied that Cr enrichment in stainless steels would result in a protective oxide layer associated with corrosion resistance against the bio-oil. Post exposure characterization showed continuous and uniform Cr distribution in the surface oxide layer of SS430F, which showed a minimal mass change, but no oxide layer on a SS430 sample, which exhibited a significant mass loss.

09 BIOMASS FUELS↗

Proceedings for the Workshop on Applied Nuclear Data Activities 2025

The 2025 Workshop for Applied Nuclear Data Activities (WANDA) covered four topic areas in nuclear data: Nuclear Data and Deterrence, Nuclear Data Prioritization for Fusion, High-Assay Low-Enriched Uranium and Novel Moderators for Advanced Reactors, and Data Preservation and Data Workflows. The intention of this workshop is to connect different communities that are invested in nuclear data and have their own unique sets of needs for the purposes of sharing information, fostering collaboration in areas of shared interest, and leveraging synergistic capabilities. The attendance of federal program managers at these workshops is essential in creating awareness of the needs of their respective communities and in providing information to better guide funding investments. In each of these topical sessions, a general description of the nuclear data needs and/or capabilities was presented, along with discussions of existing capabilities that could be leveraged, potential synergistic needs or resources, and challenges that must be overcome for the application space to progress. There are many synergistic nuclear data needs among these application spaces. The discussion largely focused on increasing the accuracy of the nuclear data and better quantifying the data uncertainties that have the greatest impact on applications. The full-day session on data processing and data workflows was by nature intended to be synergistic and applicable to all technical sessions at WANDA. A notable common theme that was highlighted across all sessions was the need for accelerated delivery of nuclear data products across complex and time-consuming workflows.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

metagRoot: a comprehensive database of protein families associated with plant root microbiomes

The plant root microbiome is vital in plant health, nutrient uptake, and environmental resilience. To explore and harness this diversity, we present metagRoot, a specialized and enriched database focused on the protein families of the plant root microbiome. MetagRoot integrates metagenomic, metatranscriptomic, and reference genome-derived protein data to characterize 71 091 enriched protein families, each containing at least 100 sequences. These families are annotated with multiple sequence alignments, CRISPR elements, hidden Markov models, taxonomic and functional classifications, ecosystem and geolocation metadata, and predicted 3D structures using AlphaFold2. MetagRoot is a powerful tool for decoding the molecular landscape of root-associated microbial communities and advancing microbiome-informed agricultural practices by enriching protein family information with ecological and structural context. The database is available at https://pavlopoulos-lab.org/metagroot/ or https://www.metagroot.org.

Chasapi, Maria N↗

MONTI: A Multi-Omics Non-negative Tensor Decomposition Framework for Gene-Level Integrative Analysis

Multi-omics data is frequently measured to enrich the comprehension of biological mechanisms underlying certain phenotypes. However, due to the complex relations and high dimension of multi-omics data, it is difficult to associate omics features to certain biological traits of interest. For example, the clinically valuable breast cancer subtypes are well-defined at the molecular level, but are poorly classified using gene expression data. Here, we propose a multi-omics analysis method called MONTI (Multi-Omics Non-negative Tensor decomposition for Integrative analysis), which goal is to select multi-omics features that are able to represent trait specific characteristics. Here, we demonstrate the strength of multi-omics integrated analysis in terms of cancer subtyping. The multi-omics data are first integrated in a biologically meaningful manner to form a three dimensional tensor, which is then decomposed using a non-negative tensor decomposition method. From the result, MONTI selects highly informative subtype specific multi-omics features. MONTI was applied to three case studies of 597 breast cancer, 314 colon cancer, and 305 stomach cancer cohorts. For all the case studies, we found that the subtype classification accuracy significantly improved when utilizing all available multi-omics data. MONTI was able to detect subtype specific gene sets that showed to be strongly regulated by certain omics, from which correlation between omics types could be inferred. Furthermore, various clinical attributes of nine cancer types were analyzed using MONTI, which showed that some clinical attributes could be well explained using multi-omics data. We demonstrated that integrating multi-omics data in a gene centric manner improves detecting cancer subtype specific features and other clinical features, which may be used to further understand the molecular characteristics of interest. The software and data used in this study are available at: https://github.com/inukj/MONTI.

59 BASIC BIOLOGICAL SCIENCES↗