Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data enhancement”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Enhancing Data Quality Monitoring at CMS with Interactive Visualization Tools and Automated Reference Run Selection

Current data quality monitoring (DQM) tools at CMS offer granularity limited to per-run analysis. Consequently, issues manifesting at the per-lumisection level can go unnoticed or, even if detectable, often lead to the classification of the whole run as bad, resulting in unnecessary data loss. Additionally, shifters have to evaluate a large set of monitoring elements during their long shifts, increasing the probability of human errors or overlooked problems. In this contribution, we present ongoing work on the development of tools that will provide shifters with an accessible, granularity-enhanced view of DQM data through interactive and dynamic visualizations. Furthermore, we introduce a reference run selection tool currently under development, which will automate the selection based on data-taking conditions and will offer a curated set of training data for machine learning models that will be used for the partial automation of the offline data certification process. These endeavors will be integrated into the DIALS website, enabling enhancements in data certification accuracy and improving the accessibility of DQM at CMS.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

A Physical Model Enhanced Data Driven Method for High-Resolution Residential Load Profile Generation

Residential buildings account for significant energy consumption, creating opportunities to offer grid services. As electric utilities seek to implement effective system operation strategies, understanding residential energy consumption patterns becomes essential; However, the time intervals of load profiles measured by utilities' smart meters are typically from 15 minutes to 60 minutes. The low-resolution data make it hard to extract appliance-level load information, which is critical for providing grid services. This paper presents a load profile generator designed to produce synthetic load profiles for residential buildings that emphasizes the importance of accurate representations of realistic energy consumption patterns. The generator takes realistic low-resolution residential load measurements and weather data as inputs, producing 1-minute interval profiles that match the characteristics of the original profiles. Further, this generator can be used to populate load profiles in areas where actual measurements are limited to improve the ability of utilities to analyze their distribution systems. By providing more high-resolution residential building load profiles, this tool supports electric utilities to enhance their residential building load control strategies and improve overall grid stability.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Phosphoproteomics Modifications in Women with Rheumatoid Arthritis─Application of Web-Based Software to Enhance Data Visualization

Individuals with rheumatoid arthritis (RA) are at increased risk of functional disability, cardiovascular disease, and obesity, all of which are influenced by dysregulated skeletal muscle. Here, this pilot study aims to identify phosphoproteomics changes in RA skeletal muscle and visualize modifications through development of a web-based app designed to promote user-friendly data interpretation and visualization. NanoLC–MS/MS analysis was performed on vastus lateralis biopsies from three women with RA and matched healthy controls. Differential analysis was performed using the Limma R package. Kinase substrate enrichment analysis (KSEA) predicted changes in kinase activity. RA muscle displayed 35 upregulated and 60 downregulated phosphosites, including the cytoskeletal proteins TTN (Ser33201, Ser33013, Ser20925), NEB (Ser2219, Thr254, Ser33013, Ser20925), FLNA (Ser1459), and LASP1 (Ser146). Compared to healthy controls, KSEA predicted decreased activity of several kinases in RA muscle, including PRKACA and CDKs. All such changes were visualized by use of our web-based app. Overall, phosphoproteome analysis reveals signaling alterations in RA skeletal muscle linked to cytoskeletal proteins, representing candidate disease biomarkers; these modifications can be explored through use of our web-based software.

phosphoproteomics↗

An Autonomous MCP Bridge to Rucio: Enhancing Data Management Accessibility for High Energy Physics

The Rucio Data Management System [1] is an important tool used by High Energy Physics experiments, including those at Fermi National Accelerator Laboratory, to store and manage exabyte-scale scientific datasets. Despite its central role in coordinating data across globally distributed storage sites, Rucio's command line interface (CLI) presents a steep learning curve, and makes it difficult for scientists to navigate through. To solve this issue, a containerized Model Context Protocol (MCP) [2] server was built that connects Large Language Models directly to Rucio, allowing AI agents to handle data tasks by using simple, natural language rather than memorized terminal commands. The core engineering focus of this project was moving the server away from slow terminal commands that require text parsing and replacing them with a native Python Client API toolset and a planned REST API framework. Moving to the Python API handles data operations directly in memory, which helps clear up formatting errors, provides the AI with clean, structured JSON data and speeds up tool execution. To prove that the system actually works, a benchmarking pipeline was also built with various questions to test the AI across four different model configurations. The questions included finding data scopes, tracking down specific datasets, and checking replication rules. Through benchmarking, early runs showed that with raw terminal text, the model would get confused and stuck, whereas switching to the Python API to feed the AI clean, structured data yielded massive improvement. By creating an intelligent and autonomous bridge to a storage network, this project shows how AI can be implemented in scientific data management, which ultimately helps scientists at Fermilab spend less time sorting through data and more time focusing on their experiments and analysis.

Akella, Kashyap [William Rainey Harper Coll.]↗

Data fusion enhanced multi-modality wellbore integrity inspection system

A downhole multi-modality inspection system includes a first imaging device operable to generate first imaging data and a second imaging device operable to generate second imaging data. The first imaging device includes a first source operable to emit energy of a first modality, and a first detector operable to detect returning energy induced by the emitted energy of the first modality. The second imaging device includes a second source operable to emit energy of a second modality, and a second detector operable to detect returning energy induced by the emitted energy of the second modality. The system further includes a processor configured to receive the first imaging data and the second imaging data, and integrate the first imaging data with the second imaging data into an enhanced data stream. The processor correlates the first imaging data and the second imaging data to provide enhanced data for detecting potential wellbore anomalies.

Kasten, Ansas Matthias↗

Techniques for storing data to enhance recovery and detection of data corruption errors

Often there are errors when reading data from computer memory. To detect and correct these errors, there are multiple types of error correction codes. Disclosed is an error correction architecture that creates a codeword having a data portion and an error correction code portion. Swizzling rearranges the order of bits and distributes the bits among different codewords. Because the data is redistributed, a potential memory error of up to N contiguous bits, where N for example equals 2 times the number of codewords swizzled together, only affects up to, at most, two bits per swizzled codeword. This keeps the error within the error detecting capabilities of the error correction architecture. Furthermore, this can allow improved error correction and detection without requiring a change to error correcting code generators and checkers.

Mills, Peter↗

Techniques for storing data to enhance recovery and detection of data corruption errors

Often there are errors when reading data from computer memory. To detect and correct these errors, there are multiple types of error correction codes. Disclosed is an error correction architecture that creates a codeword having a data portion and an error correction code portion. Swizzling rearranges the order of bits and distributes the bits among different codewords. Because the data is redistributed, a potential memory error of up to N contiguous bits, where N for example equals 2 times the number of codewords swizzled together, only affects up to, at most, two bits per swizzled codeword. This keeps the error within the error detecting capabilities of the error correction architecture. Furthermore, this can allow improved error correction and detection without requiring a change to error correcting code generators and checkers.

Mills, Peter↗

Building Stock Models for Embodied Carbon Emissions—A Review of a Nascent Field

Building stock modeling emerges as a critical tool in the strategic reduction of embodied carbon emissions, which is pivotal in reshaping the evolving construction sector. This review provides an overall view of modern methodologies in building stock modeling, homing in on the nuances of embodied carbon analysis in construction. Examining 23 seminal papers, our study delineates two primary modeling paradigms—top-down and bottom-up—each further compartmentalized into five innovative methods. This study points out the challenges of data scarcity and computational demands, advocating for methodological advancements that promise to refine the precision of building stock models. A groundbreaking trend in recent research is the incorporation of machine learning algorithms, which have demonstrated remarkable capacity, improving stock classification accuracy by 25% and urban material quantification by 40%. Furthermore, the application of remote sensing has revolutionized data acquisition, enhancing data richness by a factor of five. This review offers a critical examination of current practices and charts a course toward an environmentally prudent future. It underscores the transformative impact of building stock modeling in driving ecological stewardship in the construction industry, positioning it as a cornerstone in the quest for sustainability and its significant contribution toward the grand vision of an eco-efficient built environment.

Hu, Ming (ORCID:0000000325831161)↗

Leveraging Large Language Models for Real-World Data Evidence: A Framework for Automated Treatment Extraction and Data Harmonization

Background: The ability to comprehensively collect treatment information from cancer patient medical records would enable studies to evaluate real-world benefits and risks tied to specific treatments. Currently, it is difficult to system- atically collect high-quality treatment information because it is often stored in unstructured text. Manually extracting and standardizing drug and regimen data is time-intensive. Recent advances in large language models (LLMs) offer a potential solution for automated extraction of structured treatment information from clinical text. Objective: This study systematically evaluates the utility of four LLMs from the Llama family for automated extraction of oncology treatment information from clinical text. This information can guide researchers using cancer registry data to provide insights into cancer care and outcomes beyond clinical trials. Methods: Four instruction-tuned Llama models with varying parameter counts (1B, 3B, 8B, and 70B) were evaluated for their ability to extract treatment information from clinical documents. A unified oncology knowledge base integrating seven major public data sources was developed to standardize and normalize extracted entities—a critical step for harmonizing data from diverse sources. Extracted treatment data were compared against expert-annotated ground truth. Model performance was assessed using accuracy metrics (Precision, Recall, F1-Score) and opera- tional feasibility metrics, including processing speed and structural compliance of the output. Results: A strong positive correlation was observed between model size and extraction accuracy. F1-score improved from 0.609 for the 1B model to 0.710 (3B), 0.807 (8B), and 0.828 (70B). While larger models demonstrated superior accuracy and compliance, they incurred higher computational costs. The modest performance difference between 8B and 70B suggests diminishing returns with increasing model size. Conclusions: LLMs represent a viable technology for automating oncology treatment extraction. The 8B-parameter model emerged as a highly effective option, balancing high accuracy and computational efficiency. Selecting an appropriate LLM for deployment in cancer registries involves a trade-off between desired accuracy and available operational resources. Harmonizing extracted entities with the oncology knowledge base facilitates standardized integration into common data models, enhancing data quality for real-world evidence analyses.

artificial intelligence↗

Leveraging Inequality-Constrained Data for Enhanced Liquidus Temperature Prediction in Nuclear Waste Glass Melts

Inequality-constrained data are frequently discarded in engineering, leading to significant information loss in data-scarce domains like glass characterization in nuclear waste vitrification. This paper presents a nonparametric censored-data regression framework based on an l1-norm optimization criterion that leverages slack variables to integrate left-, right-, and interval-constrained observations into training without distributional assumptions. Validated on synthetic data and a Physics-Informed Neural Network (PINN) for predicting liquidus temperature (TL), the method improved R2 from 0.60 to 0.89 and reduced Mean Absolute Error (MAE) by 48% (51.46 to 26.89?rC) on deterministic values. The traditional models failed to satisfy any inequality constraints while the proposed l1-norm PINN satisfies 81.25% of the constraints. The proposed framework effectively extracts actionable information from previously unusable data to enhance predictive accuracy, reduce epistemic uncertainty, and ensure physical consistency in complex industrial applications.

Garcia-Morado, Erick↗

Connecting People to Data: Enabling Data Connected Communities through Enhancements to the Geothermal Data Repository

The Department of Energy's (DOE) Geothermal Data Repository (GDR) has implemented a series of new features designed to connect people to data. These features, which are based on feedback from the GDR user community and surveys of the greater geothermal research community, are designed to improve data quality and empower members of all communities to better engage with geothermal data resources by providing universal access to data and by improving the connections between data providers, subject matter experts, and the communities of people using GDR data. This paper will explore some of the recent enhancements made to the GDR to improve data discoverability, reduce submission time, and result in better quality data submissions. These improvements include the ability for users to save a list of their favorite datasets, search for insight into geothermal datasets or data availability, or sign up to receive notifications of future updates to specific datasets. These improvements aim to enhance the overall user experience of the GDR while further connecting communities to the data they need to inform decisions, advance geothermal research, and develop innovative solutions to local energy problems.

access↗

CLAS12 remote data-stream processing using ERSAP framework

Implementing a physics data processing application is relatively straightforward with the use of current containerization technologies and container image runtime services, which are prevalent in most high-performance computing (HPC) environments. However, the process is complicated by the challenges associated with data provisioning and migration, impacting the ease of workflow migration and deployment. Transitioning from traditional file-based batch processing to data-stream processing workflows is suggested as a method to streamline these workflows. This transition not only simplifies file provisioning and migration but also significantly reduces the necessity for extensive disk space. Data-stream processing is particularly effective for real-time processing during data acquisition, thereby enhancing data quality assurance. This paper introduces the integration of the JLAB CLAS12 event reconstruction application within the ERSAP data-stream processing framework that facilitates the execution of streaming event reconstruction at a remote data center and enables the return streaming of reconstructed events to JLAB while circumventing the need for temporary data storage throughout the process.

Gyurjyan, Vardan↗