Engineering PapersSearch

SEARCH · Engineering Papers

Results for “Science Data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

2023 University of Puerto Rico at Mayagüez National Institute for Congestion Reduction Study

# 2023 University of Puerto Rico at Mayagüez National Institute for Congestion Reduction Study The 2023 University of Puerto Rico at Mayagüez National Institute for Congestion Reduction (NCIR) Study provided insight into the travel patterns and associated energy consumption of participants. Study results helped researchers identify opportunities for the development of policies that could incentivize the use of alternative modes of travel such as transit and micromobility. Such travel modes reduce congestion by reducing the miles traveled by privately owned vehicles in urban and rural areas. The [National Institute for Congestion Reduction](https://nicr.usf.edu/) provides multimodal congestion reduction strategies through real-world deployments that leverage advances in technology, big data science, and innovative transportation options to optimize the efficiency and reliability of the transportation system for all users. ## Data Collection Agency The University of Puerto Rico at Mayagüez conducted the study. ## Survey Methodology The study was conducted in Spanish. Data collection was enabled via the open-source [NREL OpenPATH platform](https://www.nrel.gov/transportation/openpath). The resulting dataset consists of partially automated travel diaries—combining sensed and surveyed data reflecting patterns of multimodal, end-to-end, individual human mobility—as well as demographic and socioeconomic information from the 17 participants. ## Survey Records, Data, and Documentation Study records include 17 participants. The total number of trips was 458 and total non-air-miles traveled was approximately 1,469.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI

Curating Carbon Storage Data for Reuse: Enabling Research and Modeling from Earth’s Surface to Subsurface

The volume of public geologic carbon storage (GCS) data resources has continued to increase in recent years as the result of an increase in funding from government, industry, and academia towards national, basin, regional and field scale studies to ensure carbon capture and storage becomes a commercially viable operation. Despite the increasing volume of data, GCS data applied towards analyses such as geologic, cost, and risk modeling continues to be multi-sourced and often disparate in nature, published across government agencies, websites, data repositories and buried in derivative reports and documents. Much of the time preparing for an analysis and derivative product development is spent collecting, aggregating, transforming and preparing input data. There have been significant efforts within the DOE National Energy Technology Laboratory’s Carbon Storage Program to optimize multi-source, multi-scale subsurface geologic data curation and aggregation to support data discovery, interoperability, and reuse. Methods include the use of artificial intelligence, machine learning, and data science techniques. This talk will discuss the workflows, best practices, and processes developed to support the aggregation and curation of data through the whole system – surface to subsurface data - that support multi-scale, multi-purpose analysis for carbon storage research.

Morkner, Paige

Optimizing Batch Crystallization with Model-based Design of Experiments

Adaptive and self-optimizing intelligent systems such as digital twins are increasingly important in science and engineering. Digital twins utilize mathematical models to provide added precision to decision-making. However, physics-informed models are challenging to build, calibrate, and validate with existing data science methods. Model-based design of experiments (MBDoE) is a popular framework for optimizing data collection to maximize parameter precision in mathematical models and digital twins. In this work, we apply MBDoE, facilitated by the open-source package Pyomo.DoE, to train and validate mathematical models for batch crystallization. We quantitatively examined the estimability of the model parameters for experiments with different cooling rates. This analysis provides a quantitative explanation for the heuristic of using multiple experiments at different cooling rates.

Lynch, Hailey

Simultaneous prediction of structural properties in epitaxially–grown GaN with quantum and conventional multi–output learning algorithms

Hundreds of GaN thin film crystal plasma–assisted molecular beam epitaxy synthesis experiment records spanning two decades were organized into a dataset correlating the growth experiment design parameters with discrete, binary determinations of crystallinity and surface morphology. Conventional data science techniques as well as both quantum and classical multi–output supervised machine learning algorithms were implemented to investigate the relationships between the operating parameter data and the structural figures of merit. Correlation coefficients, decision tree nodes, p–values, and SHAP values all support substrate temperature and gallium effusion cell conditions as being statistically significant for simultaneously influencing GaN crystallinity and surface morphology. Here, a conventional deep neural network learned best from the data, followed by a quantum–classical hybrid gradient boosting algorithm. When combined with calculations of uncertainty intervals based on VennAbers predictors, machine learning predictions of both structural properties show good agreement with results reported in published experimental literature.

36 MATERIALS SCIENCE

Data analytics for intermodal freight transportation applications

With the growth of intermodal freight transportation, it is important that transportation planners and decision-makers are knowledgeable about freight flow data to make informed decisions. This is particularly true with Intelligent Transportation Systems (ITS) offering new capabilities for intermodal freight transportation. Specifically, ITS enables access to multiple different data sources, but they have different formats, resolutions, and time scales. Thus, knowledge of data science is essential to be successful in future ITS-enabled intermodal freight transportation systems. This chapter discusses the commonly used descriptive and predictive data analytic techniques in intermodal freight transportation applications. These techniques cover the entire spectrum of univariate, bivariate, and multivariate analyses. In addition to illustrating how to apply these techniques manually, this chapter will also show how to apply them using the statistical software R. Additional exercises are provided for those who wish to apply the described techniques to more complex problems.

Huynh, Nathan

Performance Year 1 Technical Report - OPEN COG Grid: Extendable Coherent Models-Datasets for Cognitive Power Grids

The OPEN COG Grid project is a collaborative effort between LLNL, NREL, and Texas A&M University (TAMU) to develop synthetic power system datasets that (i) contain all technical information that would be available in a real system, allowing to conduct studies ranging from dynamic simulation to long term planning studies; ii) are accessible to researchers from the broader data sciences community, as oppossed to power system experts only; and (iii) This report summarizes the work conducted during the first 15 months of execution of the project. These activities encompassed: 1. Conduct a survey of existing open data sets and open source power systems simulators, their supported use cases, and accessibility (Chapter 1). 2. Define a new extensible specification for power system data, covering all parameters necessary for most computational use cases (Chapter 2). 3. Collecting real technical system data to complete missing parameters in existing open source datasets (Chapter 3). 4. Develop models that capture the behavior of emergent actors in power grids, neglected by existing datasets; aggregated residential demand response (Chapter 4) and demand response of cryptocurrency miners (Chapter 5). 5. Collect detailed spatial information on distributed energy resources, particular, solar photovoltaic facilities (Chapter 6). The following chapters provide detailed descriptions of these tasks, the assumptions taken, and their findings. In conducting these tasks, the project team produced: two (accepted) conference papers; one journal paper under submission; one draft journal paper pending submission; released one repository with the developed power system data specification, with documentation and examples; and one extended dataset for the Texas power grid under review for release. The team hopes these contributions will enhance access to power system data and remove barriers to the development of new computational techniques for power systems, particularly, those inspired by cognitive sciences.

24 POWER TRANSMISSION AND DISTRIBUTION

Superionic conduction in solid polymer electrolytes – decoupling ion transport from segmental relaxation

Solvent-free, solid polymer electrolytes (SPEs) are promising candidates for next-generation, electrochemical energy storage systems due to their potential to enhance safety and performance, enable flexible device architectures, and streamline manufacturing processes. Conventional SPEs suffer from limited ionic conductivity due to the strong coupling between ion transport and (generally slow) polymer segmental relaxation. The realization of superionic conduction in SPEs, in which ions move faster than the structural relaxation of the polymers, requires a shift in design principles to promote this type of decoupled ion motion. In this perspective, we discuss how polymer architecture, ion–ion correlations, and ion–polymer interactions can unlock superionic behavior. We highlight several key design features, such as crystallinity, bulky side groups, high molecular weight, and percolating ionic aggregation, with a focus on creating low-barrier transport pathways in various polymer systems. We also demonstrate opportunities to combine polymer chemistry and data science through high-throughput and automated screening approaches to reveal how phase behavior, ion dynamics, and ionic interactions govern transport, thereby potentially enabling data-driven discovery of superionic polymer electrolyte materials.

Yang, Mengying [Univ. of Delaware, Newark, DE (Uni

United States CMM Insights Dataset

Geo-data science applications for critical mineral analysis, development acceleration, economic impact assessment, and project efficiency. Coverage: 10 years (2014-2023), 3 geographic levels (county, state, tract), 45 features Data Categories: - census: 30 features (e.g., asian population percentage, black population percentage, citizen voting age population percentage) - ejscreen: 3 features (e.g., P_DSLPM, P_PM25, P_PWDIS) - energyexpenditure: 12 features (e.g., Housing adjustment factor, Income adjustment factor, Monthly electricity cost)

AS

Leveraging public AI tools to explore systems biology resources in mathematical modeling

Predictive mathematical modeling is an essential part of systems biology and is interconnected with information management. Systems biology information is often stored in specialized formats to facilitate data storage and analysis. These formats are not designed for easy human readability and thus require specialized software to visualize and interpret results. Therefore, comprehending modeling and underlying networks and pathways is contingent on mastering systems biology tools, which is particularly challenging for users with no or little background in data science or system biology. To address this challenge, we investigated the usage of public Artificial Intelligence (AI) tools in exploring systems biology resources in mathematical modeling. We tested public AI’s understanding of mathematics in models, related systems biology data, and the complexity of model structures. Our approach can enhance the accessibility of systems biology for non-system biologists and help them understand systems biology without a deep learning curve.

59 BASIC BIOLOGICAL SCIENCES

Unraveling the Threads of Environmental Justice in Critical Mineral Extraction: A Framework for Regionalized Life Cycle Data

As we transition to a more sustainable energy system, the extraction and processing of critical minerals becomes increasingly important. However, these industries often raise concerns about environmental and social impacts, particularly in disadvantaged communities. To address these concerns, our research focuses on regionalizing environmental life cycle data to connect it with communities affected by mineral extraction. A framework was developed for collecting life cycle background data that supports the Justice40 Toolset, a market-based approach to evaluating net benefits and costs of critical mineral material recovery pathways. A goal was to alleviate public skepticism around mineral extraction processes, including secondary and unconventional feedstocks, by highlighting both environmental and social impacts. To achieve this, computational analysis and geospatial data science techniques were employed, such as within-scale and across-scale methods and proxy dataset usage. By doing so, we were able to develop a framework for identifying and disaggregating data down to regions small enough to support Justice40 goals. Our approach not only provides valuable tools fo insight into the environmental implications of mineral extraction but also helps policymakers evaluate the social impacts on local communities. This research contributes to a more just and equitable transition to a sustainable energy system, ensuring that marginalized voices are heard in decision-making process.

Davis, Tyler [NETL Site Support Contractor, Nation

Barge Science Van IMU Data

This dataset contains high-frequency (10Hz) data from the GX5-45 IMU on the Barge Science vans. The data are all raw binary files.

17 WIND ENERGY

Toward Drilling the Perfect Geothermal Well: An International Research Coordination Network for Geothermal Drilling Optimization Supported by Deep Machine Learning and Cloud Based Data Aggregation

The EDGE project, supported by the U.S. Department of Energy Geothermal Technologies Office under award DE-EE0008793, established a data-driven framework for improving the efficiency, cost-effectiveness, and reliability of geothermal well drilling. The project focused on developing scalable data infrastructure, advanced machine learning and probabilistic models, and integrated analytics tools to support continuous drilling optimization. A central objective was to reduce geothermal drilling costs by up to seventy percent while minimizing the risk of well failure through predictive diagnostics and adaptive planning. Over the project period, a comprehensive data repository was designed and deployed, incorporating records from over one hundred geothermal wells across varied geological settings. This repository supported both structured and unstructured data and adhered to FAIR data principles, enabling provenance tracking, quality control, and standardized metadata. The project introduced automated ingestion pipelines and a cloud-hosted platform that facilitated access to raw, processed, and derived datasets. This infrastructure served as the foundation for model development and analysis. Machine learning workflows were developed to predict key drilling metrics including rate of penetration, non-productive time, and total drilling costs. Self-organizing maps and dimensionality reduction methods were used to uncover operational patterns and outliers, while supervised learning algorithms such as random forests and deep neural networks were applied to forecast performance outcomes. The models were validated on heterogeneous datasets from both U.S. and Icelandic fields, demonstrating variable but significant predictive accuracy. The results indicated that finer temporal resolution, inclusion of lithological data, and consistency in operational annotations could substantially improve model performance. The project also implemented process mining techniques to reconstruct state-transition models from drilling event logs. These models enabled the identification of deviations from optimal workflows and provided insights into recurring failure modes. Analysis of non-productive time highlighted the impact of equipment failures, geological challenges, and human factors, offering opportunities for targeted mitigation strategies. The EDGE Dashboard was developed as a web-based expert system integrating data visualization, model outputs, and user-driven queries. It provided an accessible interface for operators to explore historical data, evaluate predicted outcomes, and compare drilling scenarios. Initial feedback from project partners suggested that the dashboard could serve as a foundation for more advanced advisory and optimization tools. Overall, the EDGE project demonstrated the feasibility and value of applying modern data science techniques to geothermal drilling. It delivered a set of interoperable tools and models that can support more efficient, lower-risk well development. The findings point toward a viable path for transitioning from advisory analytics to semi-autonomous drilling systems, contingent on continued collaboration, expanded datasets, and field validation. The project results have immediate relevance for drilling operations, data management practices, and future geothermal R&D efforts aimed at achieving reliable, cost-competitive geothermal energy at scale.

15 GEOTHERMAL ENERGY

Optimization of X-ray event screening using ground and in-orbit data for the Resolve instrument onboard the XRISM satellite

The X-Ray Imaging and Spectroscopy Mission (XRISM) satellite was successfully launched and put into a low-Earth orbit on September 6, 2023 (UT). The Resolve instrument onboard XRISM hosts an X-ray microcalorimeter detector, which was designed to achieve a high-resolution ( ≤ 7 eV FWHM at 6 keV), high-throughput, and non-dispersive spectroscopy over a wide energy range. It also excels in a low background with a requirement of < 2 × 10 -3 s -1 keV -1 (0.3 to 12.0 keV), which is equivalent to only one background event per spectral bin per 100-ks exposure. Event screening to discriminate X-ray events from background is a key to meeting the requirement. We present the result of the Resolve event screening using data sets recorded on the ground and in orbit based on the heritage of the preceding X-ray microcalorimeter missions, in particular, the Soft X-ray Spectrometer onboard ASTRO-H. We optimize and evaluate 19 screening items of three types based on (1) the event pulse shape, (2) relative arrival times among multiple events, and (3) good time intervals. We show that the initial screening, which is applied for science data products in the performance verification phase, reduces the background rate to 1.8 × 10 -3 s -1 keV -1 meeting the requirement. We further evaluate the additional screening utilizing the correlation among some pulse shape properties of X-ray events and show that it further reduces the background rate, particularly in the < 2 keV band. Over 0.3 to 12 keV, the background rate becomes 1.0 × 10 -3 s -1 keV -1 .

47 OTHER INSTRUMENTATION

CAML: Commutative Algebra Machine Learning─A Case Study on Protein–Ligand Binding Affinity Prediction

Recently, Suwayyid and Wei introduced commutative algebra as an emerging paradigm for machine learning and data science. In this work, we propose commutative algebra machine learning (CAML) for the prediction of protein−ligand binding affinities. Specifically, we apply persistent Stanley−Reisner theory, a key concept in combinatorial commutative algebra, to the affinity predictions of protein−ligand binding and metalloprotein−ligand binding. We present three new algorithms, i.e., element-specific commutative algebra, category-specific commutative algebra, and commutative algebra on bipartite complexes, to tackle the complexity of data involved in (metallo) protein−ligand complexes. We show that the proposed CAML outperforms other state-of-theart methods in (metallo) protein−ligand binding affinity predictions, indicating the great potential of commutative algebra learning.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Enhanced Preparation for Intelligent Cybermanufacturing Systems (EPICS)

Opportunities exist for realizing transformative advances in productivity and reductions in energy footprint through ubiquitous sensing in manufacturing environments. Enhanced Preparation for Intelligent Cybermanufacturing Systems (EPICS) is a 21-month (4 academic semesters, plus one summer) experience for graduate students that focuses on scaling the knowledge, understanding and leadership skills in the cyber manufacturing area. Masters students (8/year, 32 total) complete 2-year projects on industrially-driven project topics, rotating to internships in summer semester to work on scoping and implementation at project partners. Students complete academic training in embedded systems, process modeling, data science, and cloud-based systems design. Their projects are targeted toward sensor retrofit, process monitoring, root cause analysis, and sensor fusion.

Advanced Manufacturing

Applying the FAIR Principles to computational workflows

Recent trends within computational and data sciences show an increasing recognition and adoption of computational workflows as tools for productivity and reproducibility that also democratize access to platforms and processing know-how. As digital objects to be shared, discovered, and reused, computational workflows benefit from the FAIR principles, which stand for Findable, Accessible, Interoperable, and Reusable. The Workflows Community Initiative’s FAIR Workflows Working Group (WCI-FW), a global and open community of researchers and developers working with computational workflows across disciplines and domains, has systematically addressed the application of both FAIR data and software principles to computational workflows. We present recommendations with commentary that reflects our discussions and justifies our choices and adaptations. These are offered to workflow users and authors, workflow management system developers, and providers of workflow services as guidelines for adoption and fodder for discussion. The FAIR recommendations for workflows that we propose in this paper will maximize their value as research assets and facilitate their adoption by the wider community.

97 MATHEMATICS AND COMPUTING

Faraday Slidedeck

Faraday is a data science and visualization platform for electrochemical impedance spectroscopy. The slide deck is a visual guide with high-level information pertaining to the background, theory, and development of the application.

data warehouse

AI-assisted detector design for the EIC (AID(2)E)

Artificial Intelligence is poised to transform the design of complex, large-scale detectors like ePIC at the future Electron Ion Collider. Featuring a central detector with additional detecting systems in the far forward and far backward regions, the ePIC experiment incorporates numerous design parameters and objectives, including performance, physics reach, and cost, constrained by mechanical and geometric limits. This project aims to develop a scalable, distributed AI-assisted detector design for the EIC (AID(2)E), employing state-of-the-art multiobjective optimization to tackle complex designs. Supported by the ePIC software stack and using G EANT 4 simulations, our approach benefits from transparent parameterization and advanced AI features. The workflow leverages the PanDA and iDDS systems, used in major experiments such as ATLAS at CERN LHC, the Rubin Observatory, and sPHENIX at RHIC, to manage the compute intensive demands of ePIC detector simulations. Tailored enhancements to the PanDA system focus on usability, scalability, automation, and monitoring. Ultimately, this project aims to establish a robust design capability, apply a distributed AI-assisted workflow to the ePIC detector, and extend its applications to the design of the second detector (Detector-2) in the EIC, as well as to calibration and alignment tasks. Additionally, we are developing advanced data science tools to efficiently navigate the complex, multidimensional trade-offs identified through this optimization process.

97 MATHEMATICS AND COMPUTING