Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “database improvement”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

An Experiment with LLMs as Database Design Tutors: Persistent Equity and Fairness Challenges in Online Learning

As large language models (LLMs) continue to evolve, their capacity to replace humans as their surrogates is also improving. As increasing numbers of intelligent tutoring systems (ITSs) are embracing the integration of LLMs for digital tutoring, questions are arising as to how effective they are and if their hallucinatory behaviors diminish their perceived advantages. One critical question that is seldom asked if the availability, plurality, and relative weaknesses in the reasoning process of LLMs are contributing to the much discussed digital divide and equity and fairness in online learning. In this paper, we present an experiment with database design theory assignments and demonstrate that while their capacity to reason logically is improving, LLMs are still prone to serious errors. We demonstrate that in online learning and in the absence of a human instructor, LLMs could introduce inequity in the form of “wrongful” tutoring that could be devastatingly harmful for learners, which we call ignorant bias, in increasingly popular digital learning. We also show that significant challenges remain for STEM subjects, especially for subjects for which sound and free online tutoring systems exist. Based on the set of use cases, we formulate a possible direction for an effective ITS for online database learning classes of the future.

Jamil, Hasan M. (ORCID:0000000231243780)↗

SLIA Reference Architecture Models

The SLIA Reference Architecture Models project, sponsored by the DOE CESER Energy CyberSense Program (Oct 2024–Sep 2025), advanced LLNL’s PySCES simulation tool to better support CyTRICS Prioritization and Initial Risk Assessment (PIRA) reference architectures. Key achievements include enhancements to the PySCES transmission substation facility model, expanded asset coverage, and enhancements to the PySCES code base. Software improvements reduced code complexity, migrated PySCES to Python version 3.11, introduced an object-oriented design, and added a schema database for easier updates and validation. New features support device criticality assessments and a more precise parametric simulation mode. Remaining gaps include model validation, workflow limitations, Monte Carlo convergence issues, full device criticality metric implementation, model fidelity, and general software improvements. Continued development is recommended to address these gaps and fully align PySCES with CyTRICS PIRA requirements.

97 MATHEMATICS AND COMPUTING↗

Machine-Learning-Based Mapping and Modeling of Solar Energy with Ultra-High Spatiotemporal Granularity

Despite the rapid growth of solar energy, we still lack a dynamic, high-fidelity database that tracks the spatiotemporal variations of solar PVs and their associated infrastructures across different places at a spatially resolved scale. The absence of such data presents a barrier to various applications such as solar PV growth projection, solar energy integration, solar incentive design, and climate risk assessment. In this project, we aim to bridge this gap by developing AI-based algorithms to extract granular information about solar PV installations and their associated infrastructures (i.e., distribution grids) from widely available unstructured data like remote sensing images and street views. As a result, we have built the Solar Energy Atlas, a fine-grained, large-scale geospatial overlay of distributed solar PVs and distribution grids. On top of it, we have advanced the understanding of solar adoption and distribution grid vulnerability to climate-induced extremes. Our major contributions can be summarized as follow: (1) By developing new AI algorithms, we have built the most comprehensive solar PV spatiotemporal database covering the entire US. This is the first time we obtained the exact GPS locations, size, subtype, and installation year information for rooftop solar PVs across the US. This database can be used for solar PV growth projection, solar energy integration, solar energy policy analysis and design, and spatially-resolved climate risk assessment. (2) Leveraging this database, we have uncovered the socioeconomic driving factors that are correlated with earlier onset of solar adoption and higher saturated adoption levels. We have identified the heterogeneity in the effects of different types of financial incentives on solar adoption and provided implications for tailoring incentive design based on local income levels to promote equitable solar adoption. (3) We have developed a distribution grid GIS mapping algorithm which can obtain granular geospatial and topology information about distribution grids using multi-modal open data, reducing the dependency on hard-to-obtain smart meter data of conventional approaches. It shows effectiveness in both the U.S. and Sub-Saharan Africa. Using this algorithm, we have uncovered the non-uniform vulnerability of distribution grids to wildfires in California in the aspects of undergrounding protection and Distributed Energy Resources (DER) preparedness. This has provided important implications for improving the affordability and equity of grid adaptation approaches. (3) We have made our produced database publicly available and provided user-friendly interface to enable various stakeholders and the general public to interact with the data. We have also integrated the produced data into the Data Commons platform to enable the public to access the data and correlate it with other location-specific characteristics simply using natural language as queries. The impact of our project is three-fold: (1) New algorithms for mapping solar PVs and distribution grids across space and time, which are open source to facilitate researchers and industry; (2) New databases of solar PVs and distribution grids that have been made publicly available for engineering, social, and policy applications; (3) New understandings and actionable insights on the potential approaches to promoting solar adoption and reducing energy infrastructure vulnerabilities. In this report, we start by discussing the project background and motivation (section 5), followed by the overview of project objectives (section 6). Results and discussion for each task are presented in section 7. Significant accomplishments are summarized in section 8. This report will be concluded by discussing the paths forwards (section 9), products (section 10), and team roles (section 11).

14 SOLAR ENERGY↗

First-Principles Evaluation of Proton Hopping in Tetrahedral Oxide Motifs

Proton-conducting oxides (PCOs) are important materials used as ionic conductors for energy conversion technologies. Existing research efforts on PCO optimization and discovery generally focus on complex perovskite-based oxides that require doping and alloying to engineer oxygen deficiency and high proton conductivity. However, the variety of chemical compositions and coordination environments in oxides poses challenges for efficient materials design. In this computational study, we construct a database of simplified motifs to elucidate the relationship between fundamental materials chemistry and proton kinetics. Specifically, we focus on the zincblende crystal structure as a proxy for tetrahedral metal–oxide (M–O) coordination environments. We systematically quantified the effects of cation type, oxidation states, and M–O bond lengths on the proton hopping barrier, and found that strong M–O bonds and metal cations with large and variable oxidation states (e.g., Mo 6+ , V 5+ ) lead to smaller proton hopping barriers. By mapping the candidate cations and their preferred bond geometries onto materials databases such as the Inorganic Crystal Structure Database (ICSD) and Materials Project, we identified real materials containing the corresponding metal–oxide units. In general, we observed good agreement between the calculated proton hopping barriers obtained in real crystal structures and those predicted by our motif database. We also discuss the limitations of our model and possible future extensions to improve its predictive capabilities. Overall, our model provides a first step for the rational design and quick screening of energy-efficient PCOs.

organic↗

Results from the last DD and DT JET campaigns in the framework of the EUROfusion Tokamak Exploitation Work Package activity

JET, the only tokamak capable of operating with deuterium–tritium (D–T) fuel (since TFTR was shutdown in 1999), has provided essential experimental data to support ITER and DEMO design and operation. Within the EUROfusion Tokamak Exploitation Work Package, JET completed its final campaigns (2022–2023), culminating in the third D–T campaign (DTE3). These experiments addressed key challenges in plasma scenarios, exhaust control, and tritium management under reactor-relevant conditions. Significant progress was achieved in demonstrating ITER-like integrated scenarios with impurity seeding, achieving partial divertor detachment and high confinement ($H_{98}(y,2)$ ≈ 0.85) at 3 MA in D–T plasmas. Advanced exhaust regimes such as quasi-continuous exhaust (QCE) and X-point radiator (XPR) were successfully achieved first in D–D and then extended to D–T operation, confirming their relevance for mixed isotope operation. Operational milestones included a new world record of 69 MJ fusion energy in tritium-rich hybrid plasmas and long-pulse H-mode operation up to 60 s, contributing with unique data to the CICLOP database. Physics studies focused on peeling-limited pedestals in support of ITER and improved understanding of edge stability and impurity screening in metallic environments. Extensive usage of the shattered pellet injector (SPI) on JET provided critical information for the design of the ITER disruption mitigation system (DMS). Real-time control systems for D/T ratio control and plasma exhaust were deployed and demonstrated in D–D and D–T, while energetic particle physics investigations unfolded the role of fast ions in turbulence suppression mechanisms. Comprehensive tritium retention studies using gas balance method, post-mortem analysis, and ITER-relevant laser induced desorption spectroscopy (LIDS) diagnostics provided essential input for tritium accountancy strategies. These results are validating the ITER operational concepts, inform DEMO design, and deliver critical experience in nuclear operation and scenario integration.

disruptions↗

Facilitating Data Collection of Maintenance Events to Populate the Hydrogen Component Reliability Database (HyCReD)

The Hydrogen Component Reliability Database (HyCReD) is a collaborative project between the National Renewable Energy Laboratory, the University of Maryland, and hydrogen stakeholders to improve safety and reliability for hydrogen facilities by implementing component reliability data taxonomies that support hydrogen infrastructure failure rate analysis. The project aims to quantify failure rates of hydrogen components through high-quality data collection and analysis on root causes and maintenance needed. HyCReD provides a common database for cataloging hydrogen component failures which exists for reliability research in many other mature industries [2]. The database fills a gap for the hydrogen community by providing a scientifically rigorous approach to quantitative risk assessment (QRA), prognostic health management (PHM), and reliability-centered maintenance (RCM) analysis. High level results will be aggregated and anonymized to protect company sensitive information; detailed results will be used to help address issues of hydrogen components. These advanced analytics will support accelerated deployment of hydrogen infrastructure by enabling better: design and safety of projects (safety codes and standards development), infrastructure reliability and cost (component failure rates, maintenance protocols), and component R&D needs (robust supply chain). A key to a successful HyCReD implementation is facilitating the ease of reporting and data quality in the database that can be used for analysis. Maintenance data was a previously identified gap in initial efforts to populate and validate the database taxonomies [3]. Collection of maintenance data will be instrumental in identifying failure modes and rates, identifying incipient component failures or reduced performance, cataloging best practices for maintenance routines and methods for prognostic health management, and quantifying the risk and effect of different failure modes. Several key priorities are identified for streamlined data collection to achieve quality and detailed failure data: Applicability, Ease of Use, Accessibility, and Information Security. The HyCReD team has now begun deployment of the database to several companies and groups that have signed non-disclosure agreements to facilitate the data collection of failures in industry hydrogen refueling station infrastructure. This paper will provide an update into the process of HyCReD deployment including the development of a coding guide for facility personnel to reference and ensure data quality and consistency from one station to another as well as implementation of contextually dependent data fields of system taxonomy and formatted entries to provide ease of use. The goal is to communicate the lessons learned from the roll-out to technicians and engineers in the field, and the addition of need for high level of security to protect all stakeholders.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

SESAME: The Los Alamos National Laboratory’s Tabulated Equation of State Database Description with Extensions for Multi-phase Representations

Modeling of the thermodynamic equation of state (EOS) of various materials has had a long storied tradition at LANL. As early as 1949 Feynman, Metropolis, and Teller published a paper presenting EOS values for some elements and a methodology for calculating the EOS at high compression [1]. Cowan and Ashkin made notable methodology improvements for compressed materials throughout the 1950’s and beyond [2]. In 1971 Jack Barnes and Jerry Rood created the SESAME database and by 1972 the database became publicly available.

36 MATERIALS SCIENCE↗

SLAB: simultaneous labeling and binding affinity prediction for protein–ligand structures

Machine learning models are often used as scoring functions to predict the binding affinity of a protein–ligand complex. These models are trained with limited amounts of data with experimentally measured binding affinity values. A large number of compounds are labeled inactive through single-concentration screens without measuring binding affinities. These inactive compounds, along with the active ones, can be used to train binary classification models, while regression models are trained using compounds with binding affinities only. However, the classification and regression tasks are often handled separately, without sharing the learned feature representations. In this paper, we propose a novel model architecture that jointly performs regression and classification objectives, aiming to maximize data utilization and improve predictive performance by leveraging two complementary tasks. In our setup, the regression yields the binding affinity, whereas the classification task yields the label as active or inactive. We demonstrate our method using PDBbind, the standard 3D structure database, as well as a dataset of flavivirus protease compounds with binding affinity data. Our experiments show that the new joint training strategy improves the accuracy of the model, increasing applicability in various practical drug screening scenarios.

Biological and medical sciences↗

Unlocking nighttime mobility: Land use and accessibility in public transit for night commuters

Night commuters are integral to urban transportation systems. Essential services such as healthcare and manufacturing rely on workers who travel at night, and reliable mobility options are crucial for them. A gap exists in understanding how land use and accessibility influence public transportation use among night commuters. This study addresses this gap by using public data to explore land use and accessibility factors that affect night commuters' public transportation use in New York State. We investigated (1) the demographic characteristics of night commuters; (2) the influence of land use and accessibility on nighttime public transportation use; and (3) potential improvements to increase public transportation use and their impact. We combined data from the National Household Travel Survey with the Smart Location Database to link home locations with land use characteristics. Using logistic regression, we found that although females are generally less likely to be night commuters, they are more likely to use public transportation. Longer commute distances are associated with higher use of public transportation. Increasing job density along fixed-guideway transit routes and improving overall job accessibility via public transportation significantly enhances public transportation use among night commuters. In conclusion, this research provides actionable insights for public transportation agencies and urban planners to support night commuters, improving access and encouraging nighttime employment.

Job accessibility↗

Database and deep-learning scalability of anharmonic phonon properties by automated brute-force first-principles calculations

Understanding the anharmonic phonon properties of crystal compounds—such as phonon lifetimes and thermal conductivities—is essential for investigating and optimizing their thermal transport behaviors. These properties also impact optical, electronic, and magnetic characteristics through interactions between phonons and other quasiparticles and fields. In this study, we develop an automated first-principles workflow to calculate anharmonic phonon properties and build a comprehensive database encompassing more than 6500 inorganic compounds. Utilizing this dataset, we train a graph neural network model to predict thermal conductivity values and spectra from structural parameters, demonstrating a scaling law in which prediction accuracy improves with increasing training data size. High-throughput screening with the model enables the identification of materials exhibiting extreme thermal conductivities—both high and low. The resulting database offers valuable insights into the anharmonic behavior of phonons, thereby accelerating the design and development of advanced functional materials.

Ohnishi, Masato [University of Tokyo (Japan); Inst↗

Evaluation and optimization of flow boiling frictional pressure drop correlations using the data from traditional and next-generation refrigerants in a micro-fin tube

This study presents an experimental evaluation and optimization of flow boiling frictional pressure drop correlations for conventional and next-generation refrigerants in a horizontal micro-fin tube, with particular emphasis on the newly emerging refrigerant blends R-454C and R-455A, for which pressure-drop data in enhanced tubes remain limited. Experiments were conducted with R-410A, R-454C, R-455A, R-134a, R-1234yf, and R-1234ze(E) in a copper micro-fin tube with an inner diameter of 8.468 mm, over mass fluxes ranging from 100 to 300 kg/(m²·s) depending on the refrigerant, and evaporation temperatures of 7, 12, and 14 °C. Frictional pressure gradients were determined from measured total pressure drops after subtracting acceleration pressure drop, and the resulting database was used to assess four existing models: Kuo and Wang (1996), Cavallini et al. (1997), Goto et al. (2001), and Diani et al. (2014). The measured frictional pressure gradient increased with vapor quality and mass flux for all refrigerants and increased further at lower evaporation temperatures, with the overall trend strongly related to liquid viscosity. Among the four correlations, the Goto et al. (2001) model provided the best overall agreement with the measured data before optimization. To further improve prediction accuracy, the Kuo and Wang (1996) and Goto et al. (2001) models were optimized using the complete experimental database. After optimization, both models reduced the overall mean absolute deviation to below 15%, while the optimized Goto et al. (2001) model maintained the best and most consistent overall performance. The results provide new pressure-drop data for next-generation refrigerants and demonstrate that parameter optimization can significantly enhance the applicability of existing micro-fin-tube correlations.

Hu, Yifeng [ORNL] (ORCID:0000000242875185)↗

A Proxy Method to Bridge LCA Data Gaps Using Automated Material Classification and Probabilistic Under-Specification

Life cycle assessments (LCAs) are essential for understanding the environmental impacts of material production. However, gaps in life cycle inventory (LCI) data for material and chemical inputs present a key challenge for LCA practitioners, especially in the early design stages. Strategies for filling in these gaps require additional time and expertise, which can hinder the LCA’s completion. This study combined automatic material classification and probabilistic under-specification to create a time-efficient method to fill material LCI data gaps. To illustrate the proposed method, proxy environmental impact distributions were generated using publicly available material LCI data classified into the ChemOnt chemical taxonomy using the open-source chemical classification software ClassyFire. Input materials with data gaps were then classified into the same taxonomy, where proxy environmental impact values could be selected from the available distributions to quickly fill in any data gaps. Although these methods were applied to classify material production processes available in the Federal LCA Commons and Ecoinvent databases, they can be applied to any LCA database. This study shows that classifying materials by their chemical structure produces taxonomies with increased granularity relative to industrial classification, improving the ability of under-specified proxy data to be used for differentiating the environmental impacts of competing designs.

biological databases↗

TIGER, A thermodynamics Equilibrium Tool for Explosives Update: Improved Solver

The thermodynamic equilibrium code known as TIGER has been in use since the 1970s and is designed to calculate the performance of energetic materials during detonation at high temperatures and pressures. The original TIGER code utilized a limited database consisting of 12 gaseous and 3 condensed constituents, which were composed of the elements carbon (C), hydrogen (H), nitrogen (N), oxygen (O), and aluminum (Al). In contrast, the more modern JCZS3 database features a significantly larger dataset, including 756 gaseous species, 189 positive and negative ions, and 496 condensed constituents derived from 62 different elements. While this expanded product species database allows for the exploration of a broader range of problems relevant to our laboratory, it also introduces challenges related to convergence, particularly when both liquid and solid species coexist under a vapor dome. In this work, we present an improved TIGER solution technique aimed at accurately determining the equilibrium state for systems that can produce a diverse array of products, including both solid and liquid condensed species. This report presents the status of the TIGER solver as of the end of FY25. Ongoing efforts are focused on enhancing the solver, and the report outlines several planned improvements. We anticipate that further developments will require additional documentation in the future.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

A Decision Support System to Compile Environmental Mitigations from Hydropower Licensing Documents

The process of deciphering, extracting, and compiling information from texts dense with domain-specific terminology and technical jargon is a challenging endeavor. It demands considerable expertise and deep knowledge in the respective field, resulting in a labor-intensive process when executed by humans. Furthermore, the task of identifying multiple class labels in extensive texts presents a challenge due to intra- and inter-reader variability, making the process time-consuming and costly.We’re introducing a user-friendly graphical interface, fortified with a BERT model-powered decision support system. This advanced system aims to augment efficiency, curtail data collection time, and sustain high precision in data acquisition. It is instrumental in deciphering and synthesizing intricate texts teeming with a spectrum of expressions, even within similar mitigation categories. Such tasks traditionally demand substantial human effort and specialized knowledge in the domain.Our system is specifically engineered for the task of extracting environmental mitigation information to promote sustainable hydropower development from licenses issued by the Federal Energy Regulatory Commission (FERC). These license documents are comprehensive, each containing over 15,000 words and requiring the identification of 135 different class labels. We anticipate that our system will boost reading speed, improve the consistency of classification outputs among readers, and contribute to the development of a robust scientific database of environmental mitigations associated with the 2,000+ non-federal hydropower facilities licensed by FERC in the United States.

Yoon, Hong-Jun [ORNL] (ORCID:0000000254505878)↗

Populus VariantDB v3.2 facilitates CRISPR and functional genomics research

The success of CRISPR genome editing studies depends critically on the precision of guide RNA (gRNA) design. Sequence polymorphisms in outcrossing tree species pose design hazards that can render CRISPR genome editing ineffective. Despite recent advances in tree genome sequencing with haplotype resolution, sequence polymorphism information remains largely inaccessible to various functional genomics research efforts. The Populus VariantDB v3.2 addresses these challenges by providing a user-friendly search engine to query sequence polymorphisms of heterozygous genomes. The database accepts short sequences, such as gRNAs and primers, as input for searching against multiple poplar genomes, including hybrids, with customizable parameters. We provide examples to showcase the utilities of VariantDB in improving the precision of gRNA or primer design. The platform-agnostic nature of the probe search design makes Populus VariantDB v3.2 a versatile tool for the rapidly evolving CRISPR field and other sequence-sensitive functional genomics applications. The database schema is expandable and can accommodate additional tree genomes to broaden its user base.

59 BASIC BIOLOGICAL SCIENCES↗

CVEVOLVE

CVEvolve is an agentic AI system for autonomous algorithm discovery for scientific data processing. It creates workflows where large language model agents freely set up and configure development environments and evaluation harnesses, develop and improve data processing algorithms with designed exploration-exploitation balancing mechanisms, log history and findings in a structured database, and run holdout testing to ensure algorithm generalizability. CVEvolve offers a zero-code interface and does not require users to provide structured data and evaluation scripts.

Cherukara, MatthewJoseph [Argonne National Laborat↗

Core plasma fueling by fast inward particle transport after hydrogen pellet injection in Wendelstein 7-X

A large database of more than 1000 individual cryogenic hydrogen pellets injected into Wendelstein 7-X for plasma fueling was analyzed to improve the understanding of the three phases of the process: the ablation, deposition and transport of the pellet material. Kilohertz-sampled electron density and temperature measurements revealed a more complex drift behavior than predicted by numerical code simulation. It could be explained by the poloidal plasma E r x B- drift rotation, which plays a significant role in stellarators, but was not previously considered in pellet injection codes like HPI2. The drift results in a fast poloidal rotation of the pellet material around the plasma core, leading to an almost homogeneous deposition over the involved flux surfaces regardless of magnetic high and low field side injection geometry. Additionally, a novel fast inward directed transport mechanism (‘FIT-effect’) was observed. The effect occurs on timescales of tens of milliseconds and cannot be explained by neoclassical transport or diffusion. It might be linked to the turbulence pinch recently found in Wendelstein 7-X. When the FIT-effect occurs, the pellet particles are rapidly transferred from the deposition flux surfaces to the plasma core, causing the plasma density profile to peak, which is beneficial for confinement in Wendelstein 7-X. The large pellet injection database was statistical analyzed with regard to pellet and plasma parameters, which delivered some starting points towards developing an understanding of the physics behind the FIT-effect. The results indicate, that plasma core fueling via pellet injection is largely independent of the injection geometry in stellarators under certain conditions, reducing the technical complexity of the injection system.

Wendelstein 7-X↗

CIGS Technology Advancement via Fundamental Modeling of Defect/Impurity Interactions (Final Technical Report)

The primary goals of the proposed work were to provide modeling tools (and the associated insight which comes along with model development) for design and optimization of CuIn x Ga 1-x Se 2 (CIGS) and CdSeTe (CST) solar cell manufacturing processes and to establish the foundation for comprehensive end-to-end predictive modeling tools to enable optimization of thin film photovoltaic technology for performance, cost, yield, and reliability. The initial focus of efforts within this project was to develop coupled process/optical/device models for CIGS PV technology and to work with Siva Power to apply that TCAD (technology computer-aided design) system to improve the efficiency and reduce manufacturing costs for CIGS solar cells. Our approach to that end was to generate an extensive database of DFT calculations and to use those calculations via statistical thermodynamics methods and Monte Carlo simulation to develop and characterize models for the behavior of native defects as well as intentional and unintentional impurities, including the redistribution of the primary components of CIGS films. Increased effort went toward coupling those models for defect behavior and composition evolution to the performance of multicrystalline CIGS solar cells via prediction of doping level and recombination lifetime as function of manufacturing process. In the second budget period, the project pivoted to developing a similar system for the CdSeTe system, focused especially on understanding the role of Se/Te alloy concentration. Execution of the project resulted in the successful development of TCAD systems for both CIGS and CdSeTe thin film PV within the Synopsys Sentaurus framework by utilizing the Alagator interface. In the first budget period of the project, we developed quantitative models for the major components of CIGS PV and implemented them within a framework that couples process, optical, and device simulation. From the insights we have gained, we identified novel opportunities for enhancing CIGS solar cell performance and have laid the groundwork to further optimize the layer structure, composition profile, and thermal cycles for substantially improved efficiency and lower manufacturing costs. For the CIGS system, process changes to achieve greater than 1% absolute enhancement in efficiency were identified, but testing of those approaches was stymied by lack of a domestic CIGS manufacturing partner after the closure of Siva Power as well as Miasole. For CdSeTe, a fully capable TCAD system only became ready to apply near the end of the project period, so substantial opportunities remain to apply those models to enhance the leading thin film PV technology.

14 SOLAR ENERGY↗