Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data enrichment”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Generation of Enrichment-Dependent Thermal Neutron Scattering Data

This work details the generation of enrichment-dependent thermal neutron scattering cross sections for several crucial uranium fuel compounds. The evaluations of the thermal scattering law (TSL) and associated cross sections for uranium dioxide (UO 2 ), uranium carbide (UC), and uranium nitride (UN) were performed using standard ab initio lattice dynamics (AILD) methods. The data for uranium metal was produced using a novel hybrid approach of molecular dynamics combined with lattice dynamics methods. 235 U enrichments of 5%, 10% (LEU+), 19.75% (HALEU), 93% (HEU), and 100% were considered, in addition to natural uranium. The enrichment-dependent masses and free atom cross sections were used in the generation of elastic and inelastic thermal neutron scattering cross sections, while the calculation of the phonon density of states (DOS) and resulting TSL considered only the natural isotopic composition of uranium. The use of an identical DOS for all enrichments is expected to have minimal impact on the final data, as the small change in uranium mass should not significantly affect lattice vibrations. The cross sections are shown to exhibit significant dependence on 235 U enrichment. The submission of this data to the National Nuclear Data Center (NNDC) for release in the ENDF/B-VIII.1 database should support the design of advanced reactor concepts.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Impact of recent ENDF nuclear data update, high initial enrichment and high burnup fuel on critical experiments applicability determination via the integral index c k for burnup credit validation

In 2012, NUREG/CR-7109 reported on the validation of burnup credit calculations involving major and minor actinides and major fission products which was investigated for pressurized and boiling water reactor (PWR and BWR) fuel enrichments up to 5 wt% 235 U and assembly-average burnups up to 60 GWd/MTU. Recently, there has been interest in increasing the maximum enrichment used in PWR fuel as high as 8 wt% 235 U and correspondingly increasing the maximum assembly-average burnups to approximately 75 GWd/MTU. These proposed increases in enrichment and burnup necessitate reinvestigation of the validation basis for k eff calculations for this expanded application space. Additionally, the 2012 study was performed by using the Evaluated Nuclear Data File (ENDF)/B-VII.0 nuclear data with the SCALE 6 covariance library, and the effects of using the newly released ENDF/B-VII.1 and ENDF/B-VIII.0 nuclear data and covariance libraries should be evaluated. In this work, published in NUREG/CR-7309 in 2025, the validation assessment was performed consistently with NUREG/CR-7109: modeling irradiated fuel assemblies in the Generic Burnup Credit (GBC)-32 cask defined in NUREG/CR-6747. The TSUNAMI-3D sequence was used to generate sensitivity data for the application model, and the data were compared with sensitivity data from select benchmark models. The integral parameter c k is the metric of similarity used in this study and is consistent with NUREG/CR-7109, where a c k value in excess of 0.8 indicates sufficient similarity for use in validation. A new set of benchmark experiments with sensitivity data has been assembled for this effort. The number of experiments with available sensitivity data is now 2,104, compared to 474 in NUREG/CR-7109. This increase was facilitated by the efforts of the Nuclear Energy Agency to generate sensitivity data for a majority of the experiments in the International Criticality Safety Benchmark Evaluation Project (ICSBEP) Handbook to supplement the data available in the Oak Ridge National Laboratory (ORNL) Verified, Archived Library of Inputs and Data (VALID). The complete set of benchmarks considered here includes experiments for low-enriched uranium (LEU), intermediate enriched uranium (IEU), and a mixture of uranium and plutonium (MIX) from the ICSBEP Handbook and VALID, as well as ORNL models of the Haut Taux de Combustion (HTC) experiments and other potentially relevant models not included in VALID. The updated similarity study shows that none of the extended burnup and higher enrichment combinations considered show a significant decrease in the number of potentially applicable experiments, meaning sufficient critical experiments exist for the validation of BUC criticality safety calculations, with initial enrichments up to 8 wt% 235 U and burnups up to 80 GWd/MTU. Additionally, both the ENDF/B-VII.1 and ENDF/B-VIII.0 nuclear data libraries can be used for validation since the number of critical experiments applicable for validation increases for most cases with the most recent nuclear data compared to the previous one. As in previous BUC validation studies, the French HTC experiments are the most similar in a majority of the application cases studied, especially from representative discharge burnups ranging from 40 to 80 GWd/MTU. In conclusion, these results match the conclusions presented in NUREG/CR-7109 regarding validation of the primary actinides in BUC analyses.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Replication Data for: Quantifying oxygen induced surface enrichment of a dilute PdAu alloy catalyst

The data underlying this published work have been made publicly available in this repository as part of the IMASC Data Management Plan. This work was supported as part of the Integrated Mesoscale Architectures for Sustainable Catalysis (IMASC), an Energy Frontier Research Center funded by the U.S. Department of Energy, Office of Science, Basic Energy Sciences under Award # DE-SC0012573.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Eev (enrich Enforce Validate) With Cpefinder

This code is designed to take an existing STIX bundle with vulnerability data and enrich it with additional potential vulnerabilities to provide further insight during threat analysis. It also acts as a launch platform for other enrichments tools. The additional tools include, WAVgraph and STIXEnforcer.

Beckman, BryanR [Idaho National Laboratory (INL), ↗

Utilizing Chamber Data for Developing and Validating Climate Change Models

Controlled environment chambers (e.g. growth chambers, SPAR chambers, or open-top chambers) are useful for measuring plant ecosystem responses to climatic variables and CO2 that affect plant water relations. However, data from chambers was found to overestimate responses of C fluxes to CO2 enrichment. Chamber data may be confounded by numerous artifacts (e.g. sidelighting, edge effects, increased temperature and VPD, etc) and this limits what can be measured accurately. Chambers can be used to measure canopy level energy balance under controlled conditions and plant transpiration responses to CO2 concentration can be elucidated. However, these measurements cannot be used directly in model development or validation. The response of stomatal conductance to CO2 will be the same as in the field, but the measured response must be recalculated in such a manner to account for differences in aerodynamic conductance, temperature and VPD between the chamber and the field.

Monje, Oscar↗

Enriching the Twitter Stream Increasing Data Mining Yield and Quality Using Machine Learning

Social media data streams are important sources of real-time and historical global information for science applications. At the NASA Goddard Earth Sciences Data and Information Services Center (GES DISC), we are exploring the Twitter data stream for its potential in augmenting the validation program of NASA Earth science missions, specifically the Global Precipitation Measurement (GPM) mission. We have implemented a tweet processing infrastructure that outputs classified precipitation tweets. Inputs are "passive" tweets, along with a smaller number of tweets from "active" participants, i.e., those knowingly contributing to our effort. The "active" tweets, presumably of higher quality, enrich the Twitter stream. "Active" sources include data scraped from other social media (e.g., public Facebook posts) and data from existing crowdsourcing programs (e.g., mPING reports). In addition, there is likely relevant precipitation information in images and documents that are the end points of links often included in tweets. Information derived from these "active" sources could then be tweeted into the Twitter stream, thus enriching its quality. The objective of our current work is to mine these tweet­ linked images and documents, using neural networks, to increase the information content and quality related to precipitation. For images, we classified them as either precipitation-related or not. For training and validation, we used images obtained via the Google custom search API. We created two models: (1) by training a simple Convolutional Neural Network and (2) by using transfer learning principles to adapt a pre-trained object recognition model. For documents, both those linked to tweets and the tweet contents, we trained Hierarchical Attention Networks to determine precipitation occurrence, type, and intensity. For training and validation, we used a keyword-filtered tweet data set labelled with ground truth data from Dark Sky (an API to retrieve weather-related labels) and the National Severe Storms Laboratory's Multi­ Radar/Multi-Sensor (MRMS) system. Our results demonstrated the efficacy of our machine learning approaches for enriching the Twitter stream, to derive information potentially useful for validation of earth science satellite data.

Albayrak, Arif↗

Ontology-Enriched Specifications Enabling Findable, Accessible, Interoperable, and Reusable Marine Metagenomic Datasets in Cyberinfrastructure Systems

Marine microbial ecology requires the systematic comparison of biogeochemical and sequence data to analyze environmental influences on the distribution and variability of microbial communities. With ever-increasing quantities of metagenomic data, there is a growing need to make datasets Findable, Accessible, Interoperable, and Reusable (FAIR) across diverse ecosystems. FAIR data is essential to developing analytical frameworks that integrate microbiological, genomic, ecological, oceanographic, and computational methods. Although community standards defining the minimal metadata required to accompany sequence data exist, they haven’t been consistently used across projects, precluding interoperability. Moreover, these data are not machine-actionable or discoverable by cyberinfrastructure systems. By making ‘omic and physicochemical datasets FAIR to machine systems, we can enable sequence data discovery and reuse based on machine-readable descriptions of environments or physicochemical gradients. In this work, we developed a novel technical specification for dataset encapsulation for the FAIR reuse of marine metagenomic and physicochemical datasets within cyberinfrastructure systems. This includes using Frictionless Data Packages enriched with terminology from environmental and life-science ontologies to annotate measured variables, their units, and the measurement devices used. This approach was implemented in Planet Microbe, a cyberinfrastructure platform and marine metagenomic web-portal. Here, we discuss the data properties built into the specification to make global ocean datasets FAIR within the Planet Microbe portal. We additionally discuss the selection of, and contributions to marine-science ontologies used within the specification. Finally, we use the system to discover data by which to answer various biological questions about environments, physicochemical gradients, and microbial communities in meta-analyses. This work represents a future direction in marine metagenomic research by proposing a specification for FAIR dataset encapsulation that, if adopted within cyberinfrastructure systems, would automate the discovery, exchange, and re-use of data needed to answer broader reaching questions than originally intended.

59 BASIC BIOLOGICAL SCIENCES↗

Search for CP violation in t$\overline{\textrm{t}}$H and tH production in multilepton channels in proton-proton collisions at $\sqrt{s}$ = 13 TeV

The charge-parity (CP) structure of the Yukawa interaction between the Higgs (H) boson and the top quark is measured in a data sample enriched in the t$\overline{t}$H and tH associated production, using 138 fb -1 of data collected in proton-proton collisions at $\sqrt{s}$ = 13 TeV by the CMS experiment at the CERN LHC. The study targets events where the H boson decays via H → WW or H → ττ and the top quarks decay via t → Wb: the W bosons decay either leptonically or hadronically, and final states characterized by the presence of at least two leptons are studied. Machine learning techniques are applied to these final states to enhance the separation of CP -even from CP -odd scenarios. Two-dimensional confidence regions are set on $κ$ t and $\widetilde{k}$t, which are respectively defined as the CP -even and CP -odd top-Higgs Yukawa coupling modifiers. No significant fractional CP -odd contributions, parameterized by the quantity |$f^{Htt}_{CP}$| are observed; the parameter is determined to be |$f^{Htt}_{CP}$| = 0.59 with an interval of (0.24, 0.81) at 68% confidence level. The results are combined with previous results covering the H → ZZ and H → γγ decay modes, yielding two- and one-dimensional confidence regions on $κ$ t and $\widetilde{k}$t, while |$f^{Htt}_{CP}$| is determined to be |$f^{Htt}_{CP}$| = 0.28 with an interval of |$f^{Htt}_{CP}$| < 0.55 at 68% confidence level, in agreement with the standard model CP -even prediction of |$f^{Htt}_{CP}$| = 0.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Measurement of the top quark mass using events with a single reconstructed top quark in pp collisions at $$ \sqrt{s} $$ = 13 TeV

A measurement of the top quark mass is performed using a data sample enriched with single top quark events produced in the t channel. The study is based on proton- proton collision data, corresponding to an integrated luminosity of 35.9 fb -1 , recorded at √s = 13 TeV by the CMS experiment at the LHC in 2016. Candidate events are selected by requiring an isolated high-momentum lepton (muon or electron) and exactly two jets, of which one is identified as originating from a bottom quark. Multivariate discriminants are designed to separate the signal from the background. Optimized thresholds are placed on the discriminant outputs to obtain an event sample with high signal purity. The top quark mass is found to be $172.13^{+0.76}_{-0.77}$ GeV, where the uncertainty includes both the statistical and systematic components, reaching sub-GeV precision for the first time in this event topology. The masses of the top quark and antiquark are also determined separately using the lepton charge in the final state, from which the mass ratio and difference are determined to be $0.9952^{+0.0079}_{-0.0104}$ and $0.83^{+1.79}_{-1.35}$ GeV, respectively. The results are consistent with CPT invariance.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

SeismoGen: Seismic Waveform Synthesis Using GAN With Application to Seismic Data Augmentation

Abstract Detecting earthquake arrivals within seismic time series can be a challenging task. Visual, human detection has long been considered the gold standard but requires intensive manual labor that scales poorly to large data sets. In recent years, automatic detection methods based on machine learning have been developed to improve the accuracy and efficiency. However, the accuracy of those methods relies on access to a sufficient amount of high‐quality labeled training data, often tens of thousands of records or more. We aim to resolve this dilemma by answering two questions: (1) provided with a limited amount of reliable labeled data, can we use them to generate additional, realistic synthetic waveform data? and (2) can we use those synthetic data to further enrich the training set through data augmentation, thereby enhancing detection algorithms? To address these questions, we use a generative adversarial network (GAN), a type of machine learning model which has shown supreme capability in generating high‐quality synthetic samples in multiple domains. Once trained, our GAN model is capable of producing realistic seismic waveforms of multiple labels (noise and event classes). Applied to real Earth seismic data sets in Oklahoma, we show that data augmentation from our GAN‐generated synthetic waveforms can be used to improve earthquake detection algorithms in instances when only small amounts of labeled training data are available.

Wang, Tiantong↗

3D Continuous Forcing Dataset from 3D Constrained Variational Analysis at SGP

The continuous 3D large-scale forcing (VARANAL3D) data set derived from 3D constrained variational analysis (3DCVA) extends the conventional constrained variational analysis method by incorporating multiple sub-columns within the analysis domain. This advancement introduces spatial variability into the large-scale forcing fields, thereby enriching the data set’s applicability. The VARANAL3D data set spans from 2004 to 2018 and covers a region of 5˚×4.5˚ domain around the ARM SGP site. The analysis domain is divided into 10×9 sub-columns with 0.5˚ resolution. The 3D large-scale forcing data provides necessary variables to drive and evaluate single-column models (SCM), cloud-resolving models (CRM) ,and large-eddy simulations (LES), as well as information for testing model sensitivity to spatial variability of the large-scale forcing data, facilitating more rigorous testing and refinement of physical processes in SCM/CRM/LES.

54 ENVIRONMENTAL SCIENCES↗

Search for pair production of heavy particles decaying to a top quark and a gluon in the lepton+jets final state in proton–proton collisions at $\sqrt{s}=13\,\text {Te}\hspace{-.08em}\text {V}$

A search is presented for the pair production of new heavy resonances, each decaying into a top quark (t) or antiquark and a gluon (g). The analysis uses data recorded with the CMS detector from proton–proton collisions at a center-of-mass energy of 13 TeV at the LHC, corresponding to an integrated luminosity of 138 fb -1 . Events with one muon or electron, multiple jets, and missing transverse momentum are selected. After using a deep neural network to enrich the data sample with signal-like events, distributions in the scalar sum of the transverse momenta of all reconstructed objects are analyzed in the search for a signal. No significant deviations from the standard model prediction are found. Upper limits at 95% confidence level are set on the product of cross section and branching fraction squared for the pair production of excited top quarks in the t* → tg decay channel. The upper limits range from 120 to 0.8 fb for a t* with spin-1/2 and from 15 to 1.0 fb for a t* with spin-3/2. These correspond to mass exclusion limits up to 1050 and 1700 GeV for spin-1/2 and spin-3/2 t* particles, respectively. These are the most stringent limits to date on the existence of t* → tg resonances.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Stratigraphic Change from Ca-Sulfate to Mg-Sulfate in the Sedimentary Bedrock of Gale Crater, Mars: Recent Results from Curiosity’s APXS

Curiosity’s APXS instrument has been quantifying sulfates over >30 km of traverse in Gale crater. The rover recently arrived at sedimentary strata where orbital data predicted hydrated Mg-sulfates that may record a change to a drier paleoenvironment. The sequence of strata is in the Marker Band Valley (MBV), the ~10-m-thick, metal-rich Marker Band (MB), and strata above the MB. Here, we report recent sulfate observations by the APXS and provide constraints on the occurrence of Mg-sulfate and the implications for paleoenvironment interpretations. In sedimentary strata below the MBV, Mg-sulfate enrichment (~5-10 wt%) is generally limited to larger diagenetic nodules (~1-3 cm). Ca is positively correlated with S at proportions consistent with Ca-sulfate addition to the bedrock matrix. S variation is thus controlled primarily by Ca-sulfate, which increases ~30% in transitional units below the MBV. The MBV contains the first evidence of Mg-sulfate enrichment in the bedrock matrix, confirmed by the detection of crystalline Mg-sulfate by CheMin. The MBV bedrock has the same overall bulk composition as the underlying Mt. Sharp gp. strata, but with an additional ~5-15 wt% Mg-sulfate. The MB has contrasting sulfate content: (1) targets with very high concentrations of MnO (1.5 wt%), FeO (47 wt%), and Zn (2.2 wt%) are depleted in S and (2) targets with lower metal content have evidence of Mg-sulfate addition. Strata above the MB have a bulk composition that is distinct from other rocks in Gale. For example, the bedrock has molar Fe/Mn (50-60) and Cr/Ti (0.4-0.7) similar to basaltic soil, but ~3X higher Zn and high Ge (50 ppm). Median SO3 above the MB (15 wt%) is higher than the MBV (14 wt%) as well as strata below the MBV (~8 wt%). S does not correlate with Ca or Mg above the MBV. MgO (~9 wt%) is higher than below the MBV (~5 wt%) and the SO3/MgO (1.7) is in the same range as the Mg-sulfate-bearing MBV, suggesting Mg-sulfate enrichment. APXS data indicate that the MBV and above the MB preserve a relatively sharp vertical transition (~5-10 m) from Ca-sulfate to Mg-sulfate in the rock matrix. The sharp contacts with the sulfate-depleted MB and the notable change in bulk composition above the MB may indicate a complex depositional and/or diagenetic history under conditions where enrichments in the highly soluble Mg-sulfates were ultimately preserved.

Jeffrey Allan Berger↗

A combinatorially complete epistatic fitness landscape in an enzyme active site

Protein engineering often targets binding pockets or active sites which are enriched in epistasis—nonadditive interactions between amino acid substitutions—and where the combined effects of multiple single substitutions are difficult to predict. Few existing sequence-fitness datasets capture epistasis at large scale, especially for enzyme catalysis, limiting the development and assessment of model-guided enzyme engineering approaches. We present here a combinatorially complete, 160,000-variant fitness landscape across four residues in the active site of an enzyme. Assaying the native reaction of a thermostable β-subunit of tryptophan synthase (TrpB) in a nonnative environment yielded a landscape characterized by significant epistasis and many local optima. These effects prevent simulated directed evolution approaches from efficiently reaching the global optimum. There is nonetheless wide variability in the effectiveness of different directed evolution approaches, which together provide experimental benchmarks for computational and machine learning workflows. The most-fit TrpB variants contain a substitution that is nearly absent in natural TrpB sequences—a result that conservation-based predictions would not capture. Thus, although fitness prediction using evolutionary data can enrich in more-active variants, these approaches struggle to identify and differentiate among the most-active variants, even for this near-native function. Overall, this work presents a large-scale testing ground for model-guided enzyme engineering and suggests that efficient navigation of epistatic fitness landscapes can be improved by advances in both machine learning and physical modeling.

biocatalysis↗

A multi-ancestry genetic study of pain intensity in 598,339 veterans

Chronic pain is a common problem, with more than one-fifth of adult Americans reporting pain daily or on most days. It adversely affects the quality of life and imposes substantial personal and economic costs. Efforts to treat chronic pain using opioids had a central role in precipitating the opioid crisis. Despite an estimated heritability of 25–50%, the genetic architecture of chronic pain is not well-characterized, in part because studies have largely been limited to samples of European ancestry. To help address this knowledge gap, we conducted a cross-ancestry meta-analysis of pain intensity in 598,339 participants in the Million Veteran Program, which identified 126 independent genetic loci, 69 of which are new. Pain intensity was genetically correlated with other pain phenotypes, level of substance use and substance use disorders, other psychiatric traits, education level and cognitive traits. Integration of the genome-wide association studies findings with functional genomics data shows enrichment for putatively causal genes (n = 142) and proteins (n = 14) expressed in brain tissues, specifically in GABAergic neurons. Drug repurposing analysis identified anticonvulsants, β-blockers and calcium-channel blockers, among other drug groups, as having potential analgesic effects. Our results provide insights into key molecular contributors to the experience of pain and highlight attractive drug targets.

59 BASIC BIOLOGICAL SCIENCES↗

Calibration of a soft secondary vertex tagger using proton-proton collisions at s = 13 TeV with the ATLAS detector

Several processes studied by the ATLAS experiment at the Large Hadron Collider produce low-momentum b-flavored hadrons in the final state. This paper describes the calibration of a dedicated tagging algorithm that identifies b-flavored hadrons outside of hadronic jets by reconstructing the soft secondary vertices originating from their decays. The calibration is based on a proton-proton collision dataset at a center-of-mass energy of 13 TeV corresponding to an integrated luminosity of 140 fb -1 . Scale factors used to correct the algorithm’s performance in simulated events are extracted for the b-tagging efficiency and the mistag rate of the algorithm using a data sample enriched in $t\overline{t}$ events. Several orthogonal measurement regions are defined, binned as a function of the multiplicities of soft secondary vertices and jets containing a b-flavored hadron in the event. The mistag rate scale factors are estimated separately for events with low and high average numbers of interactions per bunch crossing. The results, which are derived from events with low missing transverse momentum, are successfully validated in a phase space characterized by high missing transverse momentum and therefore are applicable to new physics searches carried out in either phase space regime.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗