Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data enhancement”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Building Stock Models for Embodied Carbon Emissions—A Review of a Nascent Field

Building stock modeling emerges as a critical tool in the strategic reduction of embodied carbon emissions, which is pivotal in reshaping the evolving construction sector. This review provides an overall view of modern methodologies in building stock modeling, homing in on the nuances of embodied carbon analysis in construction. Examining 23 seminal papers, our study delineates two primary modeling paradigms—top-down and bottom-up—each further compartmentalized into five innovative methods. This study points out the challenges of data scarcity and computational demands, advocating for methodological advancements that promise to refine the precision of building stock models. A groundbreaking trend in recent research is the incorporation of machine learning algorithms, which have demonstrated remarkable capacity, improving stock classification accuracy by 25% and urban material quantification by 40%. Furthermore, the application of remote sensing has revolutionized data acquisition, enhancing data richness by a factor of five. This review offers a critical examination of current practices and charts a course toward an environmentally prudent future. It underscores the transformative impact of building stock modeling in driving ecological stewardship in the construction industry, positioning it as a cornerstone in the quest for sustainability and its significant contribution toward the grand vision of an eco-efficient built environment.

Hu, Ming (ORCID:0000000325831161)↗

Leveraging Large Language Models for Real-World Data Evidence: A Framework for Automated Treatment Extraction and Data Harmonization

Background: The ability to comprehensively collect treatment information from cancer patient medical records would enable studies to evaluate real-world benefits and risks tied to specific treatments. Currently, it is difficult to system- atically collect high-quality treatment information because it is often stored in unstructured text. Manually extracting and standardizing drug and regimen data is time-intensive. Recent advances in large language models (LLMs) offer a potential solution for automated extraction of structured treatment information from clinical text. Objective: This study systematically evaluates the utility of four LLMs from the Llama family for automated extraction of oncology treatment information from clinical text. This information can guide researchers using cancer registry data to provide insights into cancer care and outcomes beyond clinical trials. Methods: Four instruction-tuned Llama models with varying parameter counts (1B, 3B, 8B, and 70B) were evaluated for their ability to extract treatment information from clinical documents. A unified oncology knowledge base integrating seven major public data sources was developed to standardize and normalize extracted entities—a critical step for harmonizing data from diverse sources. Extracted treatment data were compared against expert-annotated ground truth. Model performance was assessed using accuracy metrics (Precision, Recall, F1-Score) and opera- tional feasibility metrics, including processing speed and structural compliance of the output. Results: A strong positive correlation was observed between model size and extraction accuracy. F1-score improved from 0.609 for the 1B model to 0.710 (3B), 0.807 (8B), and 0.828 (70B). While larger models demonstrated superior accuracy and compliance, they incurred higher computational costs. The modest performance difference between 8B and 70B suggests diminishing returns with increasing model size. Conclusions: LLMs represent a viable technology for automating oncology treatment extraction. The 8B-parameter model emerged as a highly effective option, balancing high accuracy and computational efficiency. Selecting an appropriate LLM for deployment in cancer registries involves a trade-off between desired accuracy and available operational resources. Harmonizing extracted entities with the oncology knowledge base facilitates standardized integration into common data models, enhancing data quality for real-world evidence analyses.

artificial intelligence↗

Leveraging Inequality-Constrained Data for Enhanced Liquidus Temperature Prediction in Nuclear Waste Glass Melts

Inequality-constrained data are frequently discarded in engineering, leading to significant information loss in data-scarce domains like glass characterization in nuclear waste vitrification. This paper presents a nonparametric censored-data regression framework based on an l1-norm optimization criterion that leverages slack variables to integrate left-, right-, and interval-constrained observations into training without distributional assumptions. Validated on synthetic data and a Physics-Informed Neural Network (PINN) for predicting liquidus temperature (TL), the method improved R2 from 0.60 to 0.89 and reduced Mean Absolute Error (MAE) by 48% (51.46 to 26.89?rC) on deterministic values. The traditional models failed to satisfy any inequality constraints while the proposed l1-norm PINN satisfies 81.25% of the constraints. The proposed framework effectively extracts actionable information from previously unusable data to enhance predictive accuracy, reduce epistemic uncertainty, and ensure physical consistency in complex industrial applications.

Garcia-Morado, Erick↗

Guidelines for Publicly Archiving Terrestrial Model Data to Enhance Usability, Intercomparison, and Synthesis

Scientific communities are increasingly publishing data to evaluate, accredit, and build on published research. However, guidelines for curating data for publication are sparse for model-related research, limiting the usability of archived simulation data. In particular, there are no established guidelines for archiving data related to terrestrial models that simulate land processes and their coupled interactions with climate. Terrestrial modelers have a unique set of challenges when publishing data due to the diversity of scientific domains, research questions, and the types and scales of simulations. Researchers in the U.S. Department of Energy’s (DOE) projects use a variety of multiscale models to advance robust predictions of terrestrial and subsurface ecosystem processes. Here, we synthesize archiving needs for data associated with different DOE models, and provide guidelines for publishing terrestrial model data components following FAIR (Findable, Accessible, Interoperable, Reusable) principles. The guidelines recommend archiving model inputs and testing data used in final simulation runs along with associated codes, workflow scripts, and metadata in public repositories. Researchers should consider archiving model outputs if they are within the storage limits of the repository. We also provide considerations for how to bundle files into different data publications with citable digital object identifiers. Finally, we identify repository features and tools that would enable storage and reuse of model data. Given the diversity of DOE terrestrial models, these guidelines are transferable to other model types and will enable efficient reuse of simulation data for purposes such as model intercomparisons, initialization, benchmarking, synthesis, and comparisons with field observations.

58 GEOSCIENCES↗

Connecting People to Data: Enabling Data Connected Communities through Enhancements to the Geothermal Data Repository

The Department of Energy's (DOE) Geothermal Data Repository (GDR) has implemented a series of new features designed to connect people to data. These features, which are based on feedback from the GDR user community and surveys of the greater geothermal research community, are designed to improve data quality and empower members of all communities to better engage with geothermal data resources by providing universal access to data and by improving the connections between data providers, subject matter experts, and the communities of people using GDR data. This paper will explore some of the recent enhancements made to the GDR to improve data discoverability, reduce submission time, and result in better quality data submissions. These improvements include the ability for users to save a list of their favorite datasets, search for insight into geothermal datasets or data availability, or sign up to receive notifications of future updates to specific datasets. These improvements aim to enhance the overall user experience of the GDR while further connecting communities to the data they need to inform decisions, advance geothermal research, and develop innovative solutions to local energy problems.

access↗

CLAS12 remote data-stream processing using ERSAP framework

Implementing a physics data processing application is relatively straightforward with the use of current containerization technologies and container image runtime services, which are prevalent in most high-performance computing (HPC) environments. However, the process is complicated by the challenges associated with data provisioning and migration, impacting the ease of workflow migration and deployment. Transitioning from traditional file-based batch processing to data-stream processing workflows is suggested as a method to streamline these workflows. This transition not only simplifies file provisioning and migration but also significantly reduces the necessity for extensive disk space. Data-stream processing is particularly effective for real-time processing during data acquisition, thereby enhancing data quality assurance. This paper introduces the integration of the JLAB CLAS12 event reconstruction application within the ERSAP data-stream processing framework that facilitates the execution of streaming event reconstruction at a remote data center and enables the return streaming of reconstructed events to JLAB while circumventing the need for temporary data storage throughout the process.

Gyurjyan, Vardan↗

Enhanced Control, Optimization, and Integration of Distributed Energy Applications (ECO-IDEA)

With support from the U.S. Department of Energy Solar Energy Technologies Office, the National Renewable Energy Laboratory (NREL) partnered with Xcel Energy, Schneider Electric, Varentec, and Electric Power Research Institute (EPRI) to meet the goals of the Enabling Extreme Real-Time Grid Integration of Solar Energy (ENERGISE) program. This project developed and validated an innovative data-enhanced hierarchical control architecture that enables the efficient, reliable, resilient, and secure operation of future distribution systems with a high penetration of distributed energy resources like solar energy. The architecture enables a hybrid control approach where a centralized control layer is complemented by distributed control algorithms for solar inverters and autonomous control of grid edge devices. It is fully interoperable and includes all the cybersecurity aspects necessary for reliable and secure system operation. The hybrid approach can seamlessly integrate multiple voltage-regulation technologies, both at central and grid-edge levels, which enables reliable and efficient system operation in the face of unpredictable conditions. The overarching goal of the Eco-Idea project is to develop, validate, and deploy a unique and innovative Data-Enhanced Hierarchical Control (DEHC) architecture that comprehensively addresses the formidable challenges associated with proliferation of high penetration of distributed PV such as reverse power flows, transients from variability of PV systems, feeder load balancing, and voltage stability. These issues are exposing the weaknesses of existing grid operations and controls - including, but not limited to, lack of grid situational awareness, heuristic and slow-acting control actions, latency of control for emergency situations, and points of failure in communications. The proposed architecture will comprehensively resolve the deficiencies of current operational settings - where monitoring and control solutions proposed across industry and academia may not be interoperable and may not coexist in the same system - and will enable an efficient, reliable, resilient, and secure operation of future distribution systems with penetration of solar energy well beyond current limits. The DEHC architecture was developed and validated rigorously through hardware-in-loop simulations in the laboratory environment and deployed on the field.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Connecting People to Data: Enabling Data Connected Communities through Enhancements to the Geothermal Data Repository: Preprint

The Department of Energy's (DOE) Geothermal Data Repository (GDR) has implemented a series of new features designed to connect people to data. These features, which are based on feedback from the GDR user community and surveys of the greater geothermal research community, are designed to improve data quality and empower members of all communities to better engage with geothermal data resources by providing universal access to data and by improving the connections between data providers, subject matter experts, and the communities of people using GDR data. This paper will explore some of the recent enhancements made to the GDR to improve data discoverability, reduce submission time, and result in better quality data submissions. These improvements include the ability for users to save a list of their favorite datasets, search for insight into geothermal datasets or data availability, or sign up to receive notifications of future updates to specific datasets. These improvements aim to enhance the overall user experience of the GDR while further connecting communities to the data they need to inform decisions, advance geothermal research, and develop innovative solutions to local energy problems.

DOE↗

Notes on Real-Beam Ground Mapping with Monopulse Radar

The spatial awareness required of modern flight systems is facilitated by images generated by ground-mapping radar. In the forward and aft directions, synthetic aperture techniques are not viable, leaving us with enhancing real aperture radar data. Enhancing real aperture radar data with monopulse information can be achieved with any of several monopulse beam sharpening techniques. Two such algorithms are discussed.

47 OTHER INSTRUMENTATION↗

Open data sets for assessing photovoltaic system reliability

Photovoltaic (PV) systems have become a cornerstone of renewable energy strategies, particularly due to the significant reduction in solar power costs over the past decade. However, the long-term reliability of PV installations presents a persistent challenge, requiring the development of advanced monitoring and predictive maintenance strategies. A wide range of data types is used to evaluate the health of PV systems, including environmental conditions, electrical performance, and inspection imagery. These data enable methodologies such as machine learning (ML) models for lifetime prediction and computer vision techniques for defect detection. However, the acquisition of high-quality and comprehensive data is difficult, particularly in terms of long-term consistency and data variety. Publicly available data sets serve as valuable resources for addressing these challenges, but they often suffer from fragmentation and are difficult to access. This paper presents a comprehensive review of existing open-source data sets related to PV degradation, analyzing their features, functionalities, and potential applications. We categorize these data sets based on the specific aspects of PV system information they cover, such as environmental conditions, operational monitoring, image inspection and module materials, and propose relevant tools and ML models for processing them. In addition, we propose practices for future data collection and usage, while also discussing potential directions in data-driven research. Our aim is to enhance data utilization and publication among researchers and industry professionals, promoting a deeper understanding of the role of data in enhancing the performance and durability of PV systems.

14 SOLAR ENERGY↗

Data-driven enhancement of coherent structure-based models for predicting instantaneous wall turbulence

Predictions of the spatial representation of instantaneous wall-bounded flows, via coherent structure-based models, are highly sensitive to the geometry of the representative structures employed by them. In this study, we propose a methodology to extract the three-dimensional (3-D) geometry of the statistically significant eddies from multi-point wall-turbulence datasets, for direct implementation into these models to improve their predictions. The methodology is employed here for reconstructing a 3-D statistical picture of the inertial wall coherent turbulence for all canonical wall-bounded flows, across a decade of friction Reynolds number (Re T ). These structures are responsible for the Re T -dependence of the skin-friction drag and also facilitate the inner-outer interactions, making them key targets of structure-based models. The empirical analysis brings out the geometric self-similarity of the large-scale wall-coherent motions and also suggests the hairpin packet as the representative flow structure for all wall-bounded flows, thereby aligning with the framework on which the attached eddy model (AEM) is based. The same framework is extended here to also model the very-large-scaled motions, with a consideration of their differences in internal versus external flows. Implementation of the empirically-obtained geometric scalings for these large structures into the AEM is shown to enhance the instantaneous flow predictions for all three velocity components. Finally, an active flow control system driven by the same geometric scalings is conceptualized, towards favourably altering the influence of the wall coherent motions on the skin-friction drag.

42 ENGINEERING↗

Persistent global greening over the last four decades using novel long-term vegetation index data with enhanced temporal consistency

Advanced Very High-Resolution Radiometer (AVHRR) satellite observations have provided the longest global daily records from 1980s, but the remaining temporal inconsistency in vegetation index datasets has hindered reliable assessment of vegetation greenness trends. To tackle this, we generated novel global long-term Normalized Difference Vegetation Index (NDVI) and Near-Infrared Reflectance of vegetation (NIRv) datasets derived from AVHRR and Moderate Resolution Imaging Spectroradiometer (MODIS). We addressed residual temporal inconsistency through three-step post processing including cross-sensor calibration among AVHRR sensors, orbital drifting correction for AVHRR sensors, and machine learning-based harmonization between AVHRR and MODIS. After applying each processing step, we confirmed the enhanced temporal consistency in terms of detrended anomaly, trend and interannual variability of NDVI and NIRv at calibration sites. Our refined NDVI and NIRv datasets showed a persistent global greening trend over the last four decades (NDVI: 0.0008 yr -1 ; NIRv: 0.0003 yr -1 ), contrasting with those without the three processing steps that showed rapid greening trends before 2000 (NDVI: 0.0017 yr -1 ; NIRv: 0.0008 yr -1 ) and weakened greening trends after 2000 (NDVI: 0.0004 yr -1 ; NIRv: 0.0001 yr -1 ). These findings highlight the importance of minimizing temporal inconsistency in long-term vegetation index datasets, which can support more reliable trend analysis in global vegetation response to climate changes.

54 ENVIRONMENTAL SCIENCES↗

Mining Smart Meter Data to Enhance Distribution Grid Observability for Behind-the-Meter Load Control: Significantly improving system situational awareness and providing valuable insights

Distributed Energy Resources (DERs) are playing an increasingly important role in power systems. In 2023, five categories of DERs-distributed solar, electric vehicles (EVs), energy storage, residential smart thermostats, and small-scale combined heat and power-are expected to contribute about 104 GW to the U.S. summer peak (see GTM, 2018). With the increasing integration of DERs in power distribution systems, distributed load control is imperative to smooth the fluctuations that they introduce. However, a main challenge is that distribution systems lack systematic situational awareness because of their limited sensors. Furthermore, most customer-level behind-the-meter (BTM) DERs, such as rooftop photovoltaics (PVs), are being integrated into distribution systems, which complicates the system monitoring and control. Furthermore, enhanced electric grid monitoring is needed to promote renewable integration while ensuring reliability, but current approaches rely on expensive sensors.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Data for "Enhancing 2-Pyrone Synthase Efficiency by High-Throughput Mass-Spectrometric Quantification and In Vitro/In Vivo Catalytic Performance Correlation"

Engineering efficient biocatalysts is essential for metabolic engineering to produce valuable bioproducts from renewable resources. However, due to the complexity of cellular metabolic networks, it is challenging to translate success in vitro into high performance in cells. To meet such a challenge, an accurate and efficient quantification method is necessary to screen a large set of mutants from complex cell culture and a careful correlation between the catalysis parameters in vitro and performance in cells is required. In this study, we employed a mass-spectrometry based high-throughput quantitative method to screen new mutants of 2-pyrone synthase (2PS) for triacetic acid lactone (TAL) biosynthesis through directed evolution in E. coli. From the process, we discovered two mutants with the highest improvement (46 fold) in titer and the fastest kcat (44 fold) over the wild type 2PS, respectively, among those reported in the literature. A careful examination of the correlation between intracellular substrate concentration, Michaelis-Menten parameters and TAL titer for these two mutants reveals that a fast reaction rate under limiting intracellular substrate concentrations is important for in-cell biocatalysis. Such properties can be tuned by protein engineering and synthetic biology to adopt these engineered proteins for the maximum activities in different intracellular environments.

catalysis↗

Data for "Enhancing Lipid Production in Plant Cells through Automated High-Throughput Genome Engineering and Phenotyping"

Plant bioengineering is a time-consuming and labor-intensive process with no guarantee of achieving desired traits. Here, we present a fast, automated, scalable, high-throughput pipeline for plant bioengineering (FAST-PB) in maize (Zea mays) and Nicotiana benthamiana. FAST-PB enables genome editing and product characterization by integrating automated biofoundry engineering of callus and protoplast cells with single-cell matrix-assisted laser desorption/ionization mass spectrometry (MALDI-MS). We first demonstrated that FAST-PB could streamline Golden Gate cloning, with the capacity to construct 96 vectors in parallel. Using FAST-PB in protoplasts, we found that PEG2050 increased transfection efficiency by over 45%. For proof-of-concept, we established a reporter-gene-free method for CRISPR editing and phenotyping via mutation of high chlorophyll fluorescence 136. We show that diverse lipids were enhanced up to 6-fold using CRISPR activation of lipid controlling genes. In callus cells, an automated transformation platform was employed to regenerate plants with enhanced lipid traits through introducing multigene cassettes. Lastly, FAST-PB enabled high-throughput single-cell lipid profiling by integrating MALDI-MS with the biofoundry, protoplast, and callus cells, differentiating engineered and unengineered cells using single-cell lipidomics. These innovations massively increase the throughput of synthetic biology, genome editing, and metabolic engineering and change what is possible using single-cell metabolomics in plants.

AI/ML↗

Integrating Applied Energy and BER Smart Data Capabilities to Develop a DOE Data Fabric for Energy-Water R&D

Focal Area(s): 1) Data acquisition and assimilation enabled by machine learning, AI, and advanced methods including experimental/network design/optimization, unsupervised learning (including deep learning), and hardware-related efforts involving AI (e.g., edge computing). Science Challenge: DOE R&D, including DOE’s Basic Energy Research (BER)’s Environmental Systems Science Division (EESSD) program and DOE’s applied energy research (AER) programs (EERE, FE, and NE) are producers and consumers of Earth systems datasets. This white paper focuses on the first topic area from the call in relation to how crosscutting resources and innovations from DOE’s EESSD and AER can be brought to bear to mutual benefit and more efficient energy-water, Earth system data resources through improved. The overarching challenge posed by this call focuses on how DOE can directly leverage artificial intelligence (AI) to engineer a substantial (paradigm-changing) improvement in Earth System Predictability? While stemming from DOE BER’s EESSD program, this is a challenge that is faced and also being addressed by DOE’s AER programs. Over the past decade plus, FE, EERE, and NE programs have made important strides towards addressing this need. These strides are in many ways highly complementary to EESSD’s MODEX efforts. Energy water systems spanning metocean to groundwater to surface water systems all are data driven whether for basic energy or applied energy. These are remote, multi-variate, complex natural, and in many cases engineered, systems. Key needs and challenges of both EESSD and AER include developing data-focused tools to enhance data search and discovery to fill in knowledge gaps (address sparse data challenge), and rapidly transform datasets, including disparate and multi-source data. Leveraging DOE on-premise computing (HPC, exascale) infrastructure supports the computing-intensive algorithms required to execute these data acquisition and transformation processes to derive enriched knowledge and data, driving AI/ML and big data analytics for these systems. The opportunity lies in combining BER and AER efforts to provide a more robust, advanced, efficient and complete computing data fabric to address energy-water data acquisition and assimilation needs which currently pose significant impediments to AI/ML predictions and research.

54 ENVIRONMENTAL SCIENCES↗