Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “model data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14

The Dark Machines Anomaly Score Challenge: Benchmark Data and Model Independent Event Classification for the Large Hadron Collider

We describe the outcome of a data challenge conducted as part of the Dark Machines (https://www.darkmachines.org) initiative and the Les Houches 2019 workshop on Physics at TeV colliders. The challenged aims to detect signals of new physics at the Large Hadron Collider (LHC) using unsupervised machine learning algorithms. First, we propose how an anomaly score could be implemented to define model-independent signal regions in LHC searches. We define and describe a large benchmark dataset, consisting of >1 billion simulated LHC events corresponding to 10\, fb^{-1} 10 f b − 1 of proton-proton collisions at a center-of-mass energy of 13 TeV. We then review a wide range of anomaly detection and density estimation algorithms, developed in the context of the data challenge, and we measure their performance in a set of realistic analysis environments. We draw a number of useful conclusions that will aid the development of unsupervised new physics searches during the third run of the LHC, and provide our benchmark dataset for future studies at https://www.phenoMLdata.org. Code to reproduce the analysis is provided at https://github.com/bostdiek/DarkMachines-UnsupervisedChallenge.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Equation-based and data-driven modeling: Open-source software current state and future directions

Here, a review of current trends in scientific computing reveals a broad shift to open-source and higher-level programming languages such as Python and growing career opportunities over the next decade. Open-source modeling tools accelerate innovation in equation-based and data-driven applications. Significant resources have been deployed to develop data-driven tools (PyTorch, TensorFlow, Scikit-learn) from tech companies that rely on machine learning services to meet business needs while keeping the foundational tools open. Open-source equation-based tools such as Pyomo, CasADi, Gekko, and JuMP are also gaining momentum according to user community and development pace metrics. Integration of data-driven and principles-based tools is emerging. New compute hardware, productivity software, and training resources have the potential to radically accelerate progress. However, long-term support mechanisms are still necessary to sustain the momentum and maintenance of critical foundational packages.

97 MATHEMATICS AND COMPUTING↗

Hysteretic temperature sensitivity of wetland CH4 fluxes explained by substrate availability and microbial activity: Model Archive

This Modeling Archive is in support of an NGEE Arctic publication "Hysteretic temperature sensitivity of wetland CH4 fluxes explained by substrate availability and microbial activity" in the Journal Biogeosciences (https://doi.org/10.5194/bg-17-5849-2020), which includes the model data used in the publication. CH4 emissions from terrestrial systems are posited to increase, which can offset mitigation efforts and accelerate climate change. Yet, the accuracy of modeled CH4 emissions is sensitive to the prescribed CH4 production (or emission) temperature dependencies that are currently uncertain. Here, we use a comprehensive biogeochemistry model (ecosys) to investigate factors modulating CH4 production and emission rates across a permafrost thaw gradient encompassing a partly thawed bog and a fully thawed fen. We find that seasonally varying substrate availability drives lower and higher modeled methanogen biomass and activity, and thereby CH4 production, during the earlier and later periods of the thawed season, respectively. Package follows the Model data archiving guidelines with data in a *.zip file with three subfolders containing *.csv files and raw model output files; a data dictionary table (data_dictionary.csv) and two model output description tables (ecosys_plantspecies_output_notes.csv and ecosys_soil_ouput_notes.csv) to explain the format and meaning of individual output variables; and a user guide as a *.pdf. A detailed model description can be found in the supplement of (Grant, 2013). The ecosys source code is available at https://doi:10.5281/zenodo.3906642. The Next-Generation Ecosystem Experiments: Arctic (NGEE Arctic), was a research effort to reduce uncertainty in Earth System Models by developing a predictive understanding of carbon-rich Arctic ecosystems and feedbacks to climate. NGEE Arctic was supported by the Department of Energy's Office of Biological and Environmental Research. The NGEE Arctic project had two field research sites: 1) located within the Arctic polygonal tundra coastal region on the Barrow Environmental Observatory (BEO) and the North Slope near Utqiagvik (Barrow), Alaska and 2) multiple areas on the discontinuous permafrost region of the Seward Peninsula north of Nome, Alaska. Through observations, experiments, and synthesis with existing datasets, NGEE Arctic provided an enhanced knowledge base for multi-scale modeling and contributed to improved process representation at global pan-Arctic scales within the Department of Energy's Earth system Model (the Energy Exascale Earth System Model, or E3SM), and specifically within the E3SM Land Model component (ELM).

54 ENVIRONMENTAL SCIENCES↗

Large-Scale Visualization of 3D Unstructured Groundwater Model Using Cave Automated Virtual Environment

The immersive three-dimensional (3D) virtual reality (VR) visualization of groundwater models allows us to deepen our understanding of aquifer systems and provide better solutions to present groundwater-related problems, such as groundwater recharge, water quality, and sustainability. Visualization assists in accurately developing groundwater models and revealing important subsurface features, including faulting, folding, and unconformity. However, assessing model accuracy poses challenges due to the complexity of geology and groundwater systems. This research demonstrates a workflow to visualize and analyze raw 3D unstructured groundwater model data using an immersive Cave Automated Virtual Environment (CAVE). To visualize the unstructured groundwater model data, the raw dataset is converted into interactive CAVE-compatible formats utilizing a set of tools: ParaView, Blender, and Unity. This enables researchers to immerse themselves in the data, identifying influential patterns and relationships. e resulting insights can inform the development of sophisticated machine-learning models for groundwater level prediction. The CAVE’s immersive capabilities allow intuitive exploration from various perspectives, providing a more holistic understanding of the factors affecting groundwater levels. These insights are crucial to improve predictive models. The CAVE results also facilitate collaborative analysis and have potential applications in training and education. is research demonstrates the value of immersive VR tools such as the CAVE for unraveling intricacies within high-dimensional scientific data to drive real-world forecasting and modeling applications.

54 ENVIRONMENTAL SCIENCES↗

Spin-Controllable Dynamics in Defect-Engineered Carbon Nanotubes as Single Photon Emitters: Data-Driven Modeling and Computations

Quantum technologies, such as quantum computing and sensing, require efficient single-photon emission (SPE) sources that operate at room temperature in telecom wavelengths. While several materials can serve as SPE sources, no single platform meets all the criteria for efficiency, ambient operation, and scalability. Single-walled carbon nanotubes (SWCNTs) with covalently attached molecules offer a promising solution. Their SPE can be easily tuned via modifications of the SWCNT's diameter, chirality, and bonded molecules, enabling emission across near-IR to telecom wavelengths at ambient conditions. However, to fully realize the potential of SWCNTs and unlock their quantum capabilities, a deeper understanding of how structural defects from molecular adducts affect their emission and competing photoexcited processes is essential. To address this gap in our knowledge, this project combined quantum chemistry calculations with data-driven methods of cheminformatics (QSAR) and machine learning (ML). The developed computational approaches have provided several design strategies for covalent functionalization of SWCNTs to improve their optical response. The collaboration with Los Alamos National Lab (LANL) enabled direct comparison of computational and experimental data, facilitating method validation. This partnership was enhanced through access to LANL's Center for Integrated Nanotechnologies (CINT) utilizing User Facility Program and summer internships, which provided three NDSU graduate students with hands-on experience at LANL. The outcomes of this project included (1) Advancing the current stage of computational methods in accurate modeling of non-adiabatic spin-dependent photoexcited dynamics and its applicability to nanosystems consisting of thousands of atoms, realized as open-access codes linked to existing DFT-based software; (2) Establishing the relationship between the structure of adducts and SWCNTs and intrinsic excitonic and spin properties of defect states for guiding novel synthetic strategies and experimental probes of chemically functionalized SWCNTs as near-IR emitting materials; (3) Generating virtual libraries of hypothetical functionalized SWCNTs for virtual screening of their chemical structures and optical properties, leveraging new functionalities of SWCNTs; (4) Offering a unique experience for NDSU graduate students that prepared them for future scientific careers related to materials modeling and big data processing. These results were summarized in 12 published journal papers and 3 recently submitted papers. One of a key finding is that the position of defect sites on the SWCNT surface primarily drives the emission redshift (up to 100 meV), while the polarity of the defect-inducing molecules has a much smaller effect (~10 meV). However, the electron-donating or withdrawing properties of a molecule influence selecting reactivity of defect sites. These insights important for optimizing synthetic protocols for desired emissions in SWCNTs. We also revealed that the interaction between two defects at various positions on the SWCNT enhances the redshift and optical activity of states, favoring strong near-IR emission. This suggests that manipulations in defect concentrations is a promising strategy for controlling efficient emission. Mostly important, the defect position was found controllable by the spin states of photoexcited intermediates: Excited aromatic molecules form ortho defects with SWCNTs at their singlet states in the presence of oxygen, while oxygen-free conditions favor para defects via the triplet-state mechanism. Additionally, a heat-activated [2+2] cycloaddition reaction facilitates divalent defect formation with fewer bonding positions that narrows emission bands. These groundbreaking findings have been experimentally validated and significantly advance our understanding of defect chemistry in SWCNTs. Using a novel encoding technique and 3D-MoRSE descriptors, we developed highly accurate ML/QSAR models to predict both the 3D structure and optical properties of SWCNTs with chemical defects. This model enabled the creation of a virtual library of 125,556 structures, providing new insights into the relationship between SWCNT-defect structure and emission.

77 NANOSCIENCE AND NANOTECHNOLOGY↗

ZENN: A thermodynamics-inspired computational framework for heterogeneous data–driven modeling

Traditional entropy-based methods—such as cross-entropy loss in classification problems—have long been essential tools for representing the information uncertainty and physical disorder in data and for developing artificial intelligence algorithms. However, the rapid growth of data across various domains has introduced new challenges, particularly the integration of heterogeneous datasets with intrinsic disparities. To address this, we introduce a zentropy-enhanced neural network (ZENN), extending zentropy theory into the data science domain via intrinsic entropy, enabling more effective learning from heterogeneous data sources. ZENN simultaneously learns both energy and intrinsic entropy components, capturing the underlying structure of multisource data. To support this, we redesign the neural network architecture to better reflect the intrinsic properties and variability inherent in diverse datasets. We demonstrate the effectiveness of ZENN on classification tasks and energy landscape reconstructions, showing its superior generalization capabilities and robustness-particularly in predicting high-order derivatives. In image and text classification tasks, ZENN demonstrates superior generalization by introducing a learnable temperature variable that models latent multisource heterogeneity, allowing it to surpass state-of-the-art models on CIFAR-10/100, BBC News, and AG News. As a practical application in materials science, we employ ZENN to reconstruct the Helmholtz energy landscape of Fe3Pt using data generated from density functional theory and capture key material behaviors, including negative thermal expansion and the critical point in the temperature–pressure space. Overall, this work presents a zentropy-grounded framework for data-driven machine learning, positioning ZENN as a versatile and robust approach for scientific problems involving complex, heterogeneous datasets.

36 MATERIALS SCIENCE↗

Data-driven models of nonautonomous systems

Nonautonomous dynamical systems are characterized by time-dependent inputs, which complicates the discovery of predictive models describing the spatiotemporal evolution of the state variables of quantities of interest from their temporal snapshots. When dynamic mode decomposition (DMD) is used to infer a linear model, this difficulty manifests itself in the need to approximate the time-dependent Koopman operators. Our approach is to approximate the original nonautonomous system with a modified system derived via a local parameterization of the time-dependent inputs. The modified system comprises a sequence of local parametric systems, which are subsequently approximated by a parametric surrogate model using the DRIPS (dimension reduction and interpolation in parameter space) framework. The offline step of DRIPS relies on DMD to build a linear surrogate model, endowed with reduced-order bases for the observables mapped from training data. The online step interpolates on suitable manifolds to construct a sequence of iterative parametric surrogate models; the target/test parameter points on these manifolds are specified by a local parameterization of the test time-dependent inputs. Here, we use numerical experimentation to demonstrate the robustness of our method and compare its performance with that of deep neural networks.

97 MATHEMATICS AND COMPUTING↗

Development of a New Chelation Model: Bioassay Data Interpretation and Dose Assessment after Plutonium Intake via Wound and Treatment with DTPA

The administration of chelation therapy to treat significant intakes of actinides, such as plutonium, affects the actinide’s normal biokinetics. In particular, it enhances the actinide’s rate of excretion, such that the standard biokinetic models cannot be applied directly to the chelation-affected bioassay data in order to estimate the intake and assess the radiation dose. Here we propose a new chelation model that can be applied to the chelation-affected bioassay data after plutonium intake via wound and treatment with DTPA. In the proposed model, chelation is assumed to occur in the blood, liver, and parts of the skeleton. Ten datasets, consisting of measurements of 14 C-DTPA, 238 Pu, and 239 Pu involving humans given radiolabeled DTPA and humans occupationally exposed to plutonium via wound and treated with chelation therapy, were used for model development. The combined dataset consisted of daily and cumulative excretion (urine and feces), wound counts, measurements of excised tissue, blood, and post-mortem tissue analyses of liver and skeleton. The combined data were simultaneously fit using the chelation model linked with a plutonium systemic model, which was linked to an ad hoc wound model. The proposed chelation model was used for dose assessment of the wound cases used in this study.

60 APPLIED LIFE SCIENCES↗

Reliable Integration of AI Data Centers at Scale – Analysis, Modeling and Synthetic Data Generation

This report analyzes the power consumption of large dynamic digital loads using the open-source MIT supercloud and SURF datasets. With an emphasis on the MIT data, we calculate important power consumption characteristics to help system operators improve generation planning and resource allocation. We also introduce a rudimentary model for generating synthetic load profiles.

97 MATHEMATICS AND COMPUTING↗

Alaskan carbon-climate feedbacks will be weaker than inferred from short-term manipulations: Alaskan Benchmark Data and Model runs

This submission aimed to assess differences in short-term step warming manipulations and long-term chronic response to climate change in Alaskan ecosystems. Briefly, climate warming is occurring fastest at high latitudes. Based on short-term field experiments, this warming is projected to stimulate soil organic matter decomposition, and promote a positive feedback to climate change. We show here that the tightly coupled, nonlinear nature of high-latitude ecosystems implies that short-term (< 10 year) warming experiments produce emergent ecosystem carbon stock temperature sensitivities inconsistent with emergent multi-decadal responses. We first demonstrate that a well-tested mechanistic ecosystem model accurately represents observed carbon cycle and active layer depth responses to short-term summer warming in four diverse Alaskan sites. We then show that short-term warming manipulations do not capture the non-linear, long-term dynamics of vegetation, and thereby soil organic matter, that occur in response to thermal, hydrological, and nutrient transformations belowground. Our results demonstrate significant spatial heterogeneity in multi-decadal Arctic carbon cycle trajectories and argue for more mechanistic models to improve predictive capabilities.The model used in the current study is available publicly (https://github.com/jinyun1tang/ECOSYS), and the current submission contains the python/ matlab codes for analyzing output from the model (includng a readme file to explain the codes). The benchmark data, also enclosed, was collected from a range of published and publicly available sources (extracted using GRABIT: https://www.mathworks.com/matlabcentral/fileexchange/7173-grabit). These sources describe warming induced changes in tundra/ boreal ecosystems.

54 ENVIRONMENTAL SCIENCES↗

Data-driven modeling of power generation for a coal power plant under cycling

Increased penetration of renewables for power generation has negatively impacted the dynamics of conventional fossil fuel-based power plants. The power plants operating on the base load are forced to cycle, to adjust to the fluctuating power demands. This results in an inefficient operation of the coal power plants, which leads up to higher operating losses. To overcome such operational challenge associated with cycling and to develop an optimal process control, this work analyzes a set of models for predicting power generation. Moreover, the power generation is intrinsically affected by the state of the power plant components, and therefore our model development also incorporates additional power plant process variables while forecasting the power generation. We present and compare multiple state-of-the-art forecasting data-driven methods for power generation to determine the most adequate and accurate model. We also develop an interpretable attention-based transformer model to explain the importance of process variables during training and forecasting. The trained deep neural network (DNN) LSTM model has good accuracy in predicting gross power generation under various prediction horizons with/without cycling events and outperforms the other models for long-term forecasting. The DNN memory-based models show significant superiority over other state-of-the-art machine learning models for short, medium and long range predictions. The transformer-based model with attention enhances the selection of historical data for multi-horizon forecasting, and also allows to interpret the significance of internal power plant components on the power generation. This newly gained insights can be used by operation engineers to anticipate and monitor the health of power plant equipment during high cycling periods.

01 COAL, LIGNITE, AND PEAT↗

Finding the missing pieces: filling gaps that impede the translation of omics data into models

High-throughput omics technologies such as DNA sequencing have made the sequencing and computational assembly of microbial genomes recovered from the environment relatively routine. Computational inference of the protein products encoded by these genomes, and the associated biochemical functions, should enable the accurate prediction and modeling of microbial metabolism, organismal interactions, and ecosystem processes. However, a lack of scalable, probabilistic protein annotation tools limits the full potential of modeling for understanding the metabolism and biogeochemical cycles of microbial communities. Our approach to improve inference of protein annotations and metabolic models relied on learning from and emulating expert manual curation, leveraging software engineering and data science best practices to scale up the throughput and accuracy of annotations and metabolic model construction, building software to objectively evaluate different annotation strategies, and more closely linking the protein annotation and metabolic model inference process. Outcomes of this research include several improved or new computational tools, including DRAM (Distilled and Refined Annotation of Metabolism) for annotating microbial genomes with protein function and metabolic traits, CAMPER (Curated Annotations for Microbial Polyphenol Enzymes and Reactions) for annotating key polyphenol metabolisms, EC-Bench for comprehensive and unbiased benchmarking of annotation tools, and several apps available via the DOE Systems Biology Knowledgebase (KBase) for building genome-scale metabolic models. We demonstrate that these tools allow us to scalably annotate and understand thousands of genomes for microbial communities from a variety of systems and test cases, including rivers, thawing permafrost, and gut microbiomes. All of these computational tools are available as open-source software, with most broadly and easily accessible to the scientific community via KBase apps.

59 BASIC BIOLOGICAL SCIENCES↗

L-VISP: LSTM Visualization for Interpretable Symptom Prediction in Patient Cohorts

Symptom modelling in head and neck cancer is challenged by the complexity of heterogeneous patient data, leading to an interest in deep learning approaches. Although Long Short-Term Memory Networks (LSTMs) have shown great results in patient risk prediction, their low interpretability requires data modellers to collaborate with clinical experts to validate the results. We present L-VISP, a human–machine solution that uses visual analytics for LSTM modelling in clinical research. L-VISP uses custom visual encodings to make multiple LSTM variants interpretable, supporting a full range of analysis, from understanding model operations and evaluating performance to interpreting results in a clinical context. We evaluate L-VISP with data modellers and a clinical oncologist and present the takeaways from this multidisciplinary collaboration.

LSTM modeling↗

Hypothetical Nuclear Reactor Facility Modeling Simulation Data Scenario Comparison

This document evaluates a hypothetical nuclear powerplant and associated protective force personnel using modern modeling and simulation tools. The facility incorporates security early in the design to consider and integrate methods to resolve security issues and vulnerabilities via the facility’s inherent design characteristics before construction. The evaluation in this document is an example only. It is not intended to recommend or evaluate the effectiveness of existing physical security requirements or identify any method that the U.S. Nuclear Regulatory Commission staff may find acceptable for complying with existing requirements.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

A two-dimensional analytical unit cell model for redox flow battery evaluation and optimization

Cell performance optimization is important for improving the overall system efficiency of a redox flow battery. To gain better insights into key controlling factors of system efficiency, this work first proposed a theoretical model for a unit flow battery cell by extending a two-dimensional analytic model to a full battery cell. Such a model is then used for cell performance optimization after validating it with experimental and numerical modeling data. With the model results, the activation, equilibrium, and pump energy losses are identified as the dominant sources of battery energy losses. Further, a guideline for reducing these sources is also proposed. Following the guideline, the mass transport coefficient is shown as a key control factor of the equilibrium energy loss and Coulombic efficiency (CE). Approaches are then proposed to improve CE and the overall system efficiency. The mass transport is also revealed as the mechanism of pump rate optimization where an optimal pump rate significantly reduces the equilibrium energy loss. With both low equilibrium and pump energy losses, an optimal electrode porosity or specific area design can further improve a battery's system efficiency based on an optimal porosity predicted by the present model. The model also demonstrates distinct behaviors and overestimation in the system efficiency when reduced to a zero-dimensional model with neglected mass transport resistance. With the new model, the guideline, and new insights, this work provides a reliable and efficient tool for the evaluation and optimization of redox flow battery design in practical applications.

25 ENERGY STORAGE↗

Data Security Defense: Modeling and Detection of Synchrophasor Data Spoofing Attack for Grid Edge

Data security and cyberattack have become critical issues in the distributed power system where adversaries can swap the source information of sensors or even spoof and alter measurements. However, the cyber security of the power system is challenged by the unpredictability and stealth of the spoofing attacks. Here, to protect the data security at the grid edge, this paper developed a synchrophasor data spoofing attack detection framework based on the time-frequency feature extraction techniques including the short-time Fourier transform (STFT) and object detection network for real-time synchrophasor data categorization and spoofing attack localization. The proposed approach outperforms earlier work in terms of spoofing attack detection and offers a vital localization function employing distributed synchrophasor sensors.

24 POWER TRANSMISSION AND DISTRIBUTION↗