Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Data Science Model”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Advanced Data Science Model for Detecting Intelligent Malware

This study focused on developing a robust artificial intelligence (AI) model capable of detecting and characterizing advanced malware in Internet of Things (IoT) devices using network data. By analyzing network traffic with various machine learning (ML) models, our AI model can identify and characterize malicious activities to significantly improve malware detection accuracy and reliability as compared to traditional methods. The developed AI/ML model was trained using network data from IoT devices, leveraging classifiers such as Random Forest, Gradient Boosting, AdaBoost, and others to optimize detection performance. This project demonstrates a scalable framework for real-time malware detection and characterization in IoT networks, capable of identifying infected devices and facilitating the necessary steps to remove or isolate them, thereby preventing further infections. Although digital twin (DT) integration is not yet implemented in the current model, it represents a promising future enhancement. By creating a virtual replica of physical IoT devices, DT technology would allow for real-time monitoring and analysis without directly accessing operational technology, thus reducing the risk of compromising or reducing the performance of actual devices. This integration would further enhance the security of IoT ecosystems, combining AI technology to better flag and detect indications of malware-infected devices within a nuclear system environment.

42 ENGINEERING↗

Enhanced Preparation for Intelligent Cybermanufacturing Systems (EPICS)

Opportunities exist for realizing transformative advances in productivity and reductions in energy footprint through ubiquitous sensing in manufacturing environments. Enhanced Preparation for Intelligent Cybermanufacturing Systems (EPICS) is a 21-month (4 academic semesters, plus one summer) experience for graduate students that focuses on scaling the knowledge, understanding and leadership skills in the cyber manufacturing area. Masters students (8/year, 32 total) complete 2-year projects on industrially-driven project topics, rotating to internships in summer semester to work on scoping and implementation at project partners. Students complete academic training in embedded systems, process modeling, data science, and cloud-based systems design. Their projects are targeted toward sensor retrofit, process monitoring, root cause analysis, and sensor fusion.

Advanced Manufacturing↗

Exploring the Whole Set of Accurate Sparse Interpretable Models

In data science applications, there are often many models that fit the data well. This phenomenon was called the Rashomon Effect by Leo Breiman. The set of good models is called the Rashomon Set, and the goal of this project is to locate, store, and study the Rashomon sets for classes of interpretable models, including decision trees and generalized additive models.

97 MATHEMATICS AND COMPUTING↗

Toward equitable environmental exposure modeling through convergence of data, open, and citizen sciences: an example of air pollution exposure modeling amidst increasing wildfire smoke

Exposure modeling is critical in environmental epidemiology and human health but may face challenges (e.g., skewed data, unequal error, context-insensitive validation, and computational demands). Modeling decisions reflect the intended use of the models and the values that modelers prioritize. We aimed to provide a conceptual framework and machine learning (ML) modeling protocols that address these issues. With 500m-gridded hourly PM 2.5 and O 3 levels in Illinois before, during, and after the 2023 Canadian wildfire season as a motivating example, we conducted modeling experiments to evaluate modeling methods, guided by three domains we propose based on theories of science: 1) Data Diversity, leveraging open and citizen science data to enhance inclusivity, parsimony, and representativeness; 2) Equitable Accuracy, ensuring fairly distributed uncertainties across subpopulations; and 3) Sustainable Modeling, balancing accuracy with reducing computational demands to promote accessibility for under-resourced researchers. Here, we found that ML with publicly available data can achieve high accuracy. Depending on methods, performance may vary substantially, even with identical input data. Large but skewed data may reduce performance. Misuse of cross-validation protocols can underestimate prediction error; although we observed R 2 s of ∼98 %, the modeled estimates varied significantly, indicating the need for careful model validation. By using new modeling protocols including representativeness-considered training and validation data and a new loss function, we achieved high agreement between estimates and ground-based measurements (e.g., R 2 = ∼90 % for PM 2.5 ; ∼80 % for O 3 ), equally distributed errors across sociodemographic strata and urban–rural divides, and reduction in computation time—from several weeks or months to a few days.

Exposure assessment↗

The 2024 “Hacking Limnology” Workshop Series and Virtual Summit: Increasing Inclusion, Participation, and Representation in the Aquatic Sciences

The 4th Aquatic Ecosystem MOdeling Network—Junior (AEMON-J) Hacking Limnology Workshop and 5th Virtual Summit: Incorporating Data Science and Open Science in the Aquatic Sciences (DSOS) convened 15–19 July 2024. During the week, these joint communities engaged in activities at the intersection of big data, open science, modeling, remote sensing, and the aquatic sciences. The weeklong event, with over 100 aquatic science practitioners and enthusiasts, followed a similar structure to previous years, comprising three days of workshops followed by two days of the virtual summit.

54 ENVIRONMENTAL SCIENCES↗

Optimizing Batch Crystallization with Model-based Design of Experiments

Adaptive and self-optimizing intelligent systems such as digital twins are increasingly important in science and engineering. Digital twins utilize mathematical models to provide added precision to decision-making. However, physics-informed models are challenging to build, calibrate, and validate with existing data science methods. Model-based design of experiments (MBDoE) is a popular framework for optimizing data collection to maximize parameter precision in mathematical models and digital twins. In this work, we apply MBDoE, facilitated by the open-source package Pyomo.DoE, to train and validate mathematical models for batch crystallization. We quantitatively examined the estimability of the model parameters for experiments with different cooling rates. This analysis provides a quantitative explanation for the heuristic of using multiple experiments at different cooling rates.

Lynch, Hailey↗

Integrating science for water security governance

Hydrological extremes are intensifying globally, increasing the complexity of decisions required to ensure water security. Advances in hydrological science, modeling, and data systems have expanded the technical frontier of water research, yet uptake of scientific insights in policy and management decisions remains limited. This persistent science–policy gap is not primarily a failure of knowledge generation or robustness, but an institutional challenge shaped by how scientific and governance systems are organized, coordinated, and connected to support the effective use of scientific knowledge. These challenges are particularly pronounced in multi-level and transboundary water governance, where decisions span jurisdictions and require coordination across institutional and political boundaries. We synthesize research at the science–policy interface and evidence from water security initiatives to show how institutional arrangements, scientific tool development, and research practices enable or constrain the sustained use of scientific knowledge in water-security governance processes. Building on these insights, we develop ‘shared decision infrastructure’ as a framing to describe how scientific knowledge is embedded within the institutional, relational, and procedural arrangements that connect science to decision-making processes over time. We translate this framing into a practical intervention roadmap centered on institutional design, tool translation, sustained co-production, and outcome-oriented evaluation to support the integration of science into ongoing governance processes. By positioning science as shared decision infrastructure, the roadmap clarifies how researchers can design scientific efforts that support more coordinated, accountable, and adaptive water security decisions amid deepening uncertainty.

M whitney, Kristen [NASA Goddard Space Flight Cent↗

Data and scripts from: “Denoising autoencoder for reconstructing sensor observation data and predicting evapotranspiration: noisy and missing values repair and uncertainty quantification”

This data package includes data and scripts from the manuscript “Denoising autoencoder for reconstructing sensor observation data and predicting evapotranspiration: noisy and missing values repair and uncertainty quantification”.The study addressed common challenges faced in environmental sensing and modeling, including uncertain input data, missing sensor observations, and high-dimensional datasets with interrelated but redundant variables. Point-scaled meteorological and soil sensor observations were perturbed with noises and missing values, and denoising autoencoder (DAE) neural networks were developed to reconstruct the perturbed data and further predict evapotranspiration. This study concluded that (1) the reconstruction quality of each variable depends on its cross-correlation and alignment to the underlying data structure, (2) uncertainties from the models were overall stronger than those from the data corruption, and (3) there was a tradeoff between reducing bias and reducing variance when evaluating the uncertainty of the machine learning models.This package includes:(1) Four ipython scripts (.ipynb): “DAE_train.ipynb” trains and evaluates DAE neural networks, “DAE_predict.ipynb” makes predictions from the trained DAE models, “ET_train.ipynb” trains and evaluates ET prediction neural networks, and “ET_predict.ipynb” makes predictions from trained ET models.(2) One python file (.py): “methods.py” includes all user-defined functions and python codes used in the ipython scripts.(3) A “sub_models” folder that includes five trained DAE neural networks (in pytorch format, .pt), which could be used to ingest input data before being fed to the downstream ET models in ‘ET_train.ipynb” or ‘ET_predict.ipynb’.(4) Two data files (.csv). Daily meteorological, vegetation, and soil data is in “df_data.csv”, where “df_meta.csv” contains the location and time information of “df_data.csv”. Each row (index) in “df_meta.csv” corresponds to each row in “df_data.csv”. These data files are formatted to follow the data structure requirements and be directly used in the ipython scripts, and they have been shuffled chronologically to train machine learning models. The meteorological and soil data was collected using point sensors between 2019-2023 at(4.a) Three shrub-dominated field sites in East River, Colorado (named “ph1”, “ph2” and “sg5” in “df_meta.csv”, where “ph1” and “ph2” were located at PumpHouse Hillslopes, and “sg5” was at Snodgrass Mountain meadow) and(4.b) One outdoor, mesoscale, and herbaceous-dominated experiment in Berkeley, California (named “tb” in “df_meta.csv”, short for Smartsoils Testbed at Lawrence Berkeley National Lab).- See "df_data_dd.csv" and "df_meta_dd.csv" for variable descriptions and the Methods section for additional data processing steps. See "flmd.csv" and "README.txt" for brief file descriptions.- All ipython scripts and python files are written in and require PYTHON language software.

54 ENVIRONMENTAL SCIENCES↗

Educational Consortium for Energy-related Data Science & Computation in Building Engineering Programs

The project spearheaded by Pennsylvania State University aims to address the growing need for integrating energy-focused computation and data science into building engineering education. As the demand for energy-efficient building designs and operations increases, the educational sector must adapt to equip future engineers with the necessary skills. This initiative responds to this need by developing a consortium that unites multiple institutions to enhance curriculum development, dataset curation, and resource sharing, thereby ensuring students are well-prepared for the evolving energy sector. The primary goal of the project is to establish a consortium that will develop and disseminate educational materials and training programs focused on energy-related data science and computation. Key accomplishments include the creation of a beta website for resource sharing, the development of training programs and standalone modules, and the curation of datasets accessible to the public. This effort will culminate in a curriculum that incorporates advanced modeling technologies and data science skills into building engineering programs.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

The 2025 “Hacking Limnology” Workshop Series and DSOS Virtual Summit: A Half Decade of Data‐Intensive Aquatic Science

The 5th Aquatic Ecosystem MOdeling Network—Junior (AEMON-J) “Hacking Limnology” Workshop and 6th Virtual Summit: Incorporating Data Science and Open Science in the Aquatic Sciences (DSOS) convened 21–25 July 2025. As in previous years (Fig. 1; Meyer and Zwart 2020; Meyer et al. 2021b, 2021c, 2022, 2024), the virtual workshops and summit were free of charge, the content was formatted to allow for broad engagement from a globally distributed audience, and workshop materials and recordings were made available on the AEMON-J/DSOS archive (Meyer et al. 2021a). In contrast to previous years, which primarily focused on inland aquatic ecosystems, this year's workshops and summit showcased a notable plurality of ecosystem types, with workshops spanning marine, riverine, and lacustrine environments. The weeklong event brought together researchers and practitioners interested in the nexus of data science, open science, and the aquatic sciences, hosting between 47 and 65 attendees at a single time and a higher number of registrants (n = 389), who might opt to access the material asynchronously.

Meyer, Michael F. [US Geological Survey, Portland,↗

Spectroscopic Online Monitoring: Using a Multi-Track Visible Spectrometer to Facilitate a Mass Balance Study in a Simulated TALSPEAK Process

Nuclear energy is a promising low-carbon energy candidate to meet the increased demand for green energy, where the integration of fuel recycling can have significant benefits for material usage and waste reduction. Utilizing in situ monitoring tools can provide ample opportunities to better control and safeguard nuclear material recycle processes while also offering knowledge and insight into real-time solution properties. The simultaneous measurement of analytical targets in multiple process locations can enable real-time mass balance and material accountancy calculations. This is demonstrated here with a mass balance study of Nd 3+ on countercurrent aqueous/organic metal extraction within a single centrifugal contactor. The Nd 3+ concentration was simultaneously monitored at the inlets and outlets of both aqueous and organic phases using a visible absorbance detector that allowed for the simultaneous measurement of up to six locations. The Nd 3+ concentration was calculated by using chemical data science algorithms, where model training sets were collected on a single track of the detector. The discussion includes addressing the challenges of using a model collected on a single track and applying it as a model across the other tracks on the detector. Each track of the detector corresponds to one measurement location on the contactor. The difference in the integrated moles of Nd 3+ between the inlet and outlet at the end of the experiment was near zero, indicating that the mass balance of this experiment was maintained. Overall, the online spectroscopic monitoring was able to follow changing solution conditions and accurately measure the concentration of Nd 3+ in different locations within the contactor system.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Soil metagenomics umbrella narrative

Implementing accessible, authentic research experiences in introductory courses is challenging, particularly at institutions serving diverse student populations. To address this gap, we developed and deployed a Course-based Undergraduate Research Experience (CURE) focused on plant-microbe interactions in General Biology II at Northeastern Illinois University (NEIU), a minority-serving institution with a diverse student body. Students grew sugar beets (Beta vulgaris), extracted DNA from the rhizoplane, and used the Department of Energy Systems Biology Knowledgebase (KBase) for bioinformatic analysis to compare microbial relative abundance in fertilized versus unfertilized soil. Over five semesters, the CURE engaged 103 students and leveraged the intuitive KBase platform to make complex sequencing data accessible. Pre/post-course survey data revealed significant increases in student self-assessed research skills, including the ability to explain results and determine the types of data to collect. Furthermore, students reported significant gains in confidence related to experimental design and hypothesis development, alongside a strong increase in familiarity with KBase. Informal faculty feedback indicated high student engagement and appreciation for the real-world connections (e.g. food systems, agriculture, and health). This scalable, low-cost model effectively integrates data science tools into the foundational curriculum, demonstrating a potent strategy for boosting research skills and broadening participation in authentic scientific inquiry among diverse undergraduate students.

59 BASIC BIOLOGICAL SCIENCES↗

Deep learning models map rapid plant species changes from citizen science and remote sensing data

Anthropogenic habitat destruction and climate change are reshaping the geographic distribution of plants worldwide. However, we are still unable to map species shifts at high spatial, temporal, and taxonomic resolution. Here, we develop a deep learning model trained using remote sensing images from California paired with half a million citizen science observations that can map the distribution of over 2,000 plant species. Our model— Deepbiosphere— not only outperforms many common species distribution modeling approaches (AUC 0.95 vs. 0.88) but can map species at up to a few meters resolution and finely delineate plant communities with high accuracy, including the pristine and clear-cut forests of Redwood National Park. These fine-scale predictions can further be used to map the intensity of habitat fragmentation and sharp ecosystem transitions across human-altered landscapes. In addition, from frequent collections of remote sensing data, Deepbiosphere can detect the rapid effects of severe wildfire on plant community composition across a 2-y time period. These findings demonstrate that integrating public earth observations and citizen science with deep learning can pave the way toward automated systems for monitoring biodiversity change in real-time worldwide.

Gillespie, Lauren E.↗

Prediction of plant complex traits via integration of multi-omics data

The formation of complex traits is the consequence of genotype and activities at multiple molecular levels. However, connecting genotypes and these activities to complex traits remains challenging. Here, we investigate whether integrating genomic, transcriptomic, and methylomic data can improve prediction for six Arabidopsis traits. We find that transcriptome- and methylome-based models have performances comparable to those of genome-based models. However, models built for flowering time using different omics data identify different benchmark genes. Nine additional genes identified as important for flowering time from our models are experimentally validated as regulating flowering. Gene contributions to flowering time prediction are accession-dependent and distinct genes contribute to trait prediction in different genotypes. Models integrating multi-omics data perform best and reveal known and additional gene interactions, extending knowledge about existing regulatory networks underlying flowering time determination. These results demonstrate the feasibility of revealing molecular mechanisms underlying complex traits through multi-omics data integration.

59 BASIC BIOLOGICAL SCIENCES↗

Toward a microscopic picture of hadronization and multi-parton processes

This project advanced the understanding of how quarks and gluons produced in high-energy collisions transform into the hadrons observed in particle detectors, a fundamental process known as quantum chromodynamics (QCD) hadronization. By combining theoretical calculations, quantum simulation methods, and modern AI techniques, the research developed new tools to study multi-parton dynamics and nonperturbative effects that are essential for interpreting data from current and future nuclear physics experiments. Key outcomes include new theoretical frameworks for jet and hadron measurements, pioneering quantum simulation algorithms for real-time dynamics in field theories, and the development of advanced machine-learning models, such as diffusion models and explainable classifiers, to simulate and analyze collider events. These results are directly relevant to experiments at Jefferson Lab, Brookhaven National Laboratory, and the future Electron-Ion Collider, and they also have a broader impact in areas such as quantum information science and data-driven modeling of complex systems. The project supported the training of graduate students and postdoctoral fellows and contributed to the broader scientific community through publications, workshops, and collaborative activities. Overall, this work provides new insights into the microscopic mechanisms of hadron formation and establishes a foundation for future studies at the intersection of nuclear physics, artificial intelligence, and quantum computing.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Measure this, not that: Optimizing the cost and model-based information content of measurements

Model-based design of experiments (MBDoE) is a powerful framework for selecting and calibrating science-based mathematical models from data. Here, this work extends popular MBDoE workflows by proposing a convex mixed integer (non)linear programming (MINLP) to optimize the selection of measurements. The solver MindtPy is modified to support calculating the D-optimality objective and its gradient via an external package, scipy, using the grey-box module in Pyomo. The new approach is demonstrated in two case studies: estimating highly correlated kinetics from a batch reactor and estimating transport parameters in a large-scale rotary packed bed for CO 2 capture. Both case studies show how examining the Pareto optimal trade-offs between information content measured by A- and D-optimality versus measurement budget offers practical guidance for selecting measurements for scientific experiments.

97 MATHEMATICS AND COMPUTING↗

DOE FAIR Surrogate Benchmarks Supporting AI and Simulation Research (SBI Surrogate Benchmark Initiative) (Final Report)

Computational Science is being revolutionized by integrating AI and simulation and, in particular, by deep learning surrogate models that can replace all or part of traditional large‐scale HPC computations. Such surrogates can achieve remarkable performance improvements, as much as several orders of magnitude, and save both compute time and energy. The Surrogate Benchmark Initiative (SBI) project creates a community repository and FAIR (Findable, Accessible, Interoperable, and Reusable) data ecosystem for HPC application surrogate benchmarks. The SBI team comes from Argonne National Laboratory (ANL), Indiana University (IU), Rutgers University, the University of Tennessee, Knoxville (UTK), and the University of Virginia(UVA). SBI repositories include data, code, and all relevant collateral artifacts, that the science and engineering community needs to use and reuse these data sets and surrogates. SBI repositories generate active research from both participants in SBI and the broader AI and domain science communities. This project develops surrogates that use several different neural nets to learn and quickly infer the results of simulations and data systems and capture them as surrogate benchmarks with a rich set of metadata, covering. Data; Model; Metrics specification; Machine specification; Science, Speed, Power Results, We research FAIR metadata for these benchmarks. We develop application surrogate examples as benchmarks across many fields (ANL, UTK, IU, UVA). We also study non Surrogate benchmarks that have many common features and similar issues regarding FAIRness. We work with MLCommons (UVA, UTK), which is a major machine learning benchmarking activity where we get metadata ontologies, software, and benchmarks, benchmarks have datasets, models, and metadata, and they need a technical framework developed by UTK and Rutgers and deployed by UVA. We study features of Surrogates, including performance, training set size, and uncertainty quantification (Rutgers, UVA and IU).

97 MATHEMATICS AND COMPUTING↗