Engineering PapersSearch

SEARCH · Engineering Papers

Results for “data usage”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Performability modeling based on real data: A case study

Described is a measurement-based performability model based on error and resource usage data collected on a multiprocessor system. A method for identifying the model structure is introduced and the resulting model is validated against real data. Model development from the collection of raw data to the estimation of the expected reward is described. Both normal and error behavior of the system are characterized. The measured data show that the holding times in key operational and error states are not simple exponentials and that a semi-Markov process is necessary to model system behavior. A reward function, based on the service rate and the error rate in each state, is then defined in order to estimate the performability of the system and to depict the cost of apparent types of errors.

Hsueh, M. C.

A Digital Twin Framework Utilizing Machine Learning for Robust Predictive Maintenance: Enhancing Tire Health Monitoring

We introduce a novel digital twin (DT) framework for the predictive maintenance of long-term physical systems. Using monitoring tire health as an application, we show how the DT framework can be used to enhance automotive safety and efficiency, and how the technical challenges can be overcome using a three-step approach. First, to manage the data complexity over a long operation span, we employ data reduction techniques to concisely represent physical tires using historical performance and usage data. Relying on these data, for fast real-time prediction, we train a transformer-based model offline on our concise dataset to predict future tire health over time, represented as remaining casing potential (RCP). Based on our architecture, our model quantifies both epistemic and aleatoric uncertainties, providing reliable confidence intervals around predicted RCP. Second, to incorporate real-time data, we update the predictive model in the DT framework, ensuring its accuracy throughout its lifespan with the aid of hybrid modeling and the use of the discrepancy function. Third, to assist decision-making in predictive maintenance, we implement a tire state decision algorithm, which strategically determines the optimal timing for tire replacement based on RCP forecasted by our transformer model. This approach ensures that our DT accurately predicts system health, continually refines its digital representation, and supports predictive maintenance decisions. Furthermore, our framework effectively embodies a physical system, leveraging big data and machine learning (ML) for predictive maintenance, model updates, and decision-making.

advanced computing infrastructure

Learning from GES DISC's MLS and OMI Data Users: Metrics Matter

It has been over 15 years since Aura research satellite launched in 2004 to observe the Earth's ozone layer, air quality, and climate from four different instruments - the High Resolution Dynamics Limb Sounder (HIRDLS), the Microwave Limb Sounder (MLS), the Ozone Monitoring Instrument (OMI), and the Tropospheric Emission Spectrometer (TES). Observations from the Aura mission have established a concrete understanding of the changing chemistry of our atmosphere.The NASA Goddard Earth Sciences Data and Information Services Center (GES DISC) is the official archive and distribution center for the HIRDLS, MLS, and OMI instruments. This presentation will report metrics of data usage and services on these instruments. Key to GES DISC's mission to provide better data support is gaining a better understanding of our users' needs and behaviors as they discover, access and utilize these data. We will summarize the users' needs from these instruments based on user inquiry information collected over the Aura mission lifetime and present findings from this ensemble metrics.

Metrics

Energy efficient integrated photonic systems based on inverse design

The energy footprint of modern information processing and communications systems is immense and utilizes a significant fraction of total global energy usage. Data centers alone consume over 70 billion kilowatt-hours per year. Much of this energy usage is intrinsic to the use of electronic wiring, making optical-based technologies a necessary and promising route to mitigating energy consumption in short- to medium-distance communication links. In this project, we developed and implemented a framework for the design of optical components, based on machine learning, which enables optical components relevant to optical information processing to be realized at their physical performance limits.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI

Status of the FAA flight loads monitoring program

In support of the Federal Aviation Administration Aging Aircraft Research Program, the Agency has established a Flight Loads Data Collection Program for Commercial Aircraft. The objectives of this Program are to: - Review existing data collected by other sources including but not limited to U.S., Foreign, Military, etc. - Collect current operational usage data from both large and small transport aircraft - Develop criteria for future generations of transports. This paper presents the status of the various programs which are completed, underway or planned. The FAA will be collecting, storing, and analyzing the data which characterize typical commercial transport operations. The airframe manufacturers will handle the task of calculating the loads and stresses.

Aging Aircraft

The D3 Middleware Architecture

DARWIN is a NASA developed, Internet-based system for enabling aerospace researchers to securely and remotely access and collaborate on the analysis of aerospace vehicle design data, primarily the results of wind-tunnel testing and numeric (e.g., computational fluid-dynamics) model executions. DARWIN captures, stores and indexes data; manages derived knowledge (such as visualizations across multiple datasets); and provides an environment for designers to collaborate in the analysis of test results. DARWIN is an interesting application because it supports high-volumes of data. integrates multiple modalities of data display (e.g., images and data visualizations), and provides non-trivial access control mechanisms. DARWIN enables collaboration by allowing not only sharing visualizations of data, but also commentary about and views of data. Here we provide an overview of the architecture of D3, the third generation of DARWIN. Earlier versions of DARWIN were characterized by browser-based interfaces and a hodge-podge of server technologies: CGI scripts, applets, PERL, and so forth. But browsers proved difficult to control, and a proliferation of computational mechanisms proved inefficient and difficult to maintain. D3 substitutes a pure-Java approach for that medley: A Java client communicates (though RMI over HTTPS) with a Java-based application server. Code on the server accesses information from JDBC databases, distributed LDAP security services, and a collaborative information system. D3 is a three tier-architecture, but unlike 'E-commerce' applications, the data usage pattern suggests different strategies than traditional Enterprise Java Beans - we need to move volumes of related data together, considerable processing happens on the client, and the 'business logic' on the server-side is primarily data integration and collaboration. With D3, we are extending DARWIN to handle other data domains and to be a distributed system, where a single login allows a user transparent access to test results from multiple servers and authority domains.

Walton, Joan

Global Precipitation Measurement (GPM) Mission Products and Services at the NASA Goddard Earth Sciences (GES) Data and Information Services Center (DISC)

On February 27, 2014, the NASA Global Precipitation Measurement (GPM) mission was launched to provide the next-generation global observations of rain and snow (http:pmm.nasa.govGPM). The GPM mission consists of an international network of satellites in which a GPM Core Observatory satellite carries both active and passive microwave instruments to measure precipitation and serve as a reference standard, to unify precipitation measurements from a constellation of other research and operational satellites. The NASA Goddard Earth Sciences (GES) Data and Information Services Center (DISC) hosts and distributes GPM data within the NASA Earth Observation System Data Information System (EOSDIS). The GES DISC is home to the data archive for the GPM predecessor, the Tropical Rainfall Measuring Mission (TRMM). Over the past 16 years, the GES DISC has served the scientific as well as other communities with TRMM data and user-friendly services. During the GPM era, the GES DISC will continue to provide user-friendly data services and customer support to users around the world. GPM products currently and to-be available include the following: 1. Level-1 GPM Microwave Imager (GMI) and partner radiometer products. 2. Goddard Profiling Algorithm (GPROF) GMI and partner products. 3. Integrated Multi-satellitE Retrievals for GPM (IMERG) products. (early, late, and final)A dedicated Web portal (including user guides, etc.) has been developed for GPM data (http:disc.sci.gsfc.nasa.govgpm). Data services that are currently and to-be available include Google-like Mirador (http:mirador.gsfc.nasa.gov) for data search and access; data access through various Web services (e.g., OPeNDAP, GDS, WMS, WCS); conversion into various formats (e.g., netCDF, HDF, KML (for Google Earth), ASCII); exploration, visualization, and statistical online analysis through Giovanni (http:giovanni.gsfc.nasa.gov); generation of value-added products; parameter and spatial subsetting; time aggregation; regridding; data version control and provenance; documentation; science support for proper data usage, FAQ, help desk; monitoring services (e.g. Current Conditions) for applications.In this presentation, we will present GPM data products and services with examples.

NASA GES DISC

Mic-hackathon 2024: hackathon on machine learning for electron and scanning probe microscopy

Microscopy is one of the primary sources of information on materials structure and functionality at the nanometer and atomic scales. The data generated through microscopy is often contained in well-structured datasets, enriched with extensive metadata and sample histories, although not always with the same level of detail or storage format. The broad incorporation of data management plans by major funding agencies ensures the preservation and accessibility of this data. However, deriving insights from these rich datasets remains challenging due to the lack of established code ecosystems, standardized benchmarks, and integration strategies. Correspondingly, the efficiency of data usage is very low, and time expenditures at the analysis stage are enormous. In addition to post-acquisition data analysis, the emergence of application programming interfaces by major microscope manufacturers now creates opportunities for real-time ML-based data analytics to enable automated decision making, and particularly ML-agent controlled real-time microscope operation. Despite these opportunities, there is a significant gap in integrating the ML community with the broader microscopy community, limiting the value that these methods bring to physics and materials discovery and materials optimization. Hackathons address these challenges by fostering collaboration between ML experts and microscopy professionals, encouraging the development of innovative solutions that leverage ML for microscopy and preparing the workforce of the future both for microscopy-intensive domains areas, instrument manufacturers, and ML scientists interested in real world applications for fundamental research, materials optimization, and manufacturing. The hackathon generated benchmark datasets and digital twins of microscopes that further contribute to the development of the field and establish data analysis ecosystems. All the codes can be found at GitHub(https://github.com/KalininGroup/Mic-hackathon-2024-codes-publication/tree/1.0.0.1) and Zenodo (https://zenodo.org/records/15579940).

97 MATHEMATICS AND COMPUTING

Scale-up Unlearnable Examples Learning with High-performance Computing

Recent advancements in AI models, like ChatGPT, are structured to retain user interactions, which could inadvertently include sensitive healthcare data. In the healthcare field, particularly when radiologists use AI-driven diagnostic tools hosted on online platforms, there is a risk that medical imaging data may be repurposed for future AI training without explicit consent, spotlighting critical privacy and intellectual property concerns around healthcare data usage. Addressing these privacy challenges, a novel approach known as Unlearnable Examples (UEs) has been introduced, aiming to make data unlearnable to deep learning models. A prominent method within this area, called Unlearnable Clustering (UC), has shown improved UE performance with larger batch sizes but was previously limited by computational resources (e.g., a single workstation). To push the boundaries of UE performance with theoretically unlimited resources, we scaled up UC learning across various datasets using Distributed Data Parallel (DDP) training on the Summit supercomputer. Our goal was to examine UE efficacy at high-performance computing (HPC) levels to prevent unauthorized learning and enhance data security, particularly exploring the impact of batch size on UE’s unlearnability. Utilizing the robust computational capabilities of the Summit, extensive experiments were conducted on diverse datasets such as Pets, MedMNist, Flowers, and Flowers102. Our findings reveal that both overly large and overly small batch sizes can lead to performance instability and affect accuracy. However, the relationship between batch size and unlearnability varied across datasets, highlighting the necessity for tailored batch size strategies to achieve optimal data protection. The use of Summit’s high-performance GPUs, along with the efficiency of the DDP framework, facilitated rapid updates of model parameters and consistent training across nodes. Our results underscore the critical role of selecting appropriate batch sizes based on the specific characteristics of each dataset to prevent learning and ensure data security in deep learning applications. The source code is publicly available at https: // github. com/ hrlblab/ UE_ HPC .

Zhu, Yanfan [Vanderbilt University, Nashville, TN,

A Regional Perspective on Global NWP from North America and Recent Developments in the NASA GEOS System

Satellite data have played an important role in improving model forecast skills. This presentation will give a perspective of data usages of vital satellites on global NWP and show some examples of using existing satellite observations in the GEOS data assimilation system at NASA GMAO. The efforts to utilize emerging satellite data and to prepare for the upcoming new instruments NASA supports will be presented as well.

Yanqiu Zhu

Electricity use in big area additive manufacturing of fiber-reinforced polymer composites

In recent years, additive manufacturing (AM), especially large-format additive manufacturing (LFAM), has gained momentum in the manufacturing industry. While LFAM offers benefits over conventional manufacturing processes, such as minimizing material waste and providing vast geometric freedom, assessing its sustainability remains challenging due to limited data, particularly on energy consumption. Most existing data pertain to small-scale or desktop AM and are not directly applicable to LFAM. In this study, we conducted real-time measurements of electricity usage for a type of LFAM known as big area additive manufacturing (BAAM), which typically uses fiber-reinforced polymer pellets as feedstock. We collected electricity usage data from fifteen printing jobs over two months in an industrial production setting. These data fill the existing gap and can be reused to enhance the community’s understanding of LFAM electricity usage, support further research, and promote sustainable development in advanced manufacturing technologies.

ecology

Why Is Rainfall Error Analysis Requisite for Data Assimilation and Climate Modeling?

Given the large temporal and spatial variability of precipitation processes, errors in rainfall observations are difficult to quantify yet crucial to making effective use of rainfall data for improving atmospheric analysis, weather forecasting, and climate modeling. We highlight the need for developing a quantitative understanding of systematic and random errors in precipitation observations by examining explicit examples of how each type of errors can affect forecasts and analyses in global data assimilation. We characterize the error information needed from the precipitation measurement community and how it may be used to improve data usage within the general framework of analysis techniques, as well as accuracy requirements from the perspective of climate modeling and global data assimilation.

Hou, Arthur Y.

Distributed and Secure Spectrum Sharing for 5G and 6G Networks

Secure spectrum sharing or spectrum co-existence of multiple 5G networks and future 6G networks is a powerful enabler technology. The National Spectrum Strategy (NSS) published by the White House in November, 2023, and the subsequent NSS implementation plan led by the National Telecommunication and Information Administration (NTIA) is the driver of a national effort to enable co-existence of government incumbents and commercial networks in selected spectrum bands. Cellular networks such as 5G & 6G and non-cellular Wi-Fi 6E & 7 are the prominent wireless technologies considered for co-existence with incumbent wireless links. Security of the spectrum sharing solutions is a must to make this transformation of spectrum use possible, specially for mission critical communications. However, current spectrum sharing solutions rely on centralized data bases with inherent vulnerabilities. This paper focuses on secure spectrum sharing among multiple 5G networks using unlicensed and shared frequency bands. It presents an innovative AI/ML based distributed spectrum sharing approach that can be autonomously used by multiple networks. Each sharing network uses its own observation of the Radio Frequency (RF) environment, which consists of RF measurements reported from the 5G User Equipment (UE), to adjust the transmission power levels for secure co-existence. Data is presented to illustrate the superior performance of this solution compared to other spectrum sharing solutions where each network can utilize usage data of the other networks. Finally it discusses how this efficient spectrum sharing solution can evolve in the future for the 6G networks.

5G

Measurement-based reliability/performability models

Measurement-based models based on real error-data collected on a multiprocessor system are described. Model development from the raw error-data to the estimation of cumulative reward is also described. A workload/reliability model is developed based on low-level error and resource usage data collected on an IBM 3081 system during its normal operation in order to evaluate the resource usage/error/recovery process in a large mainframe system. Thus, both normal and erroneous behavior of the system are modeled. The results provide an understanding of the different types of errors and recovery processes. The measured data show that the holding times in key operational and error states are not simple exponentials and that a semi-Markov process is necessary to model the system behavior. A sensitivity analysis is performed to investigate the significance of using a semi-Markov process, as opposed to a Markov process, to model the measured system.

Hsueh, Mei-Chen

Using CALIPSO's New Ocean Derived Column Optical Depths

CALIPSO’s Version 4.51 Level 2 data release introduces an all-new group of science data sets containing estimates of total column two-way transmittances and effective optical depths derived from CALIOP ocean surface backscatter measurements and MERRA-2 reanalysis wind speed data. These estimates use data from the standard CALIOP lidar signal but in a passive sensor-like way, thus creating a unique link to passive instrument measurements. These new retrievals are provided for the entire mission, day and night, at 532nm, and are reported at single shot, 1km, and 5km resolutions for all profiles in which a valid lidar ocean surface return is detected. The addition of a total column optical depth constraint on subsequent retrievals of cloud and aerosol optical properties opens the door for many exciting new ways to leverage the already rich and versatile CALIPSO dataset. Following a brief review of the retrieval technique, this talk will focus on quality assurance assessments, estimated uncertainties, and data usage scenarios. We will conclude with examples highlighting some of the exciting work already being done using this new addition to CALIPSO’s already rich data record.

R Ryan

Estimating Flexibility Envelopes for Residential Customers From Utility Smart Meter Data: Preprint

Demand response from residential customers has significant potential to support power system operations, but accurate flexibility estimation is challenging due to the limited resolution of advanced metering infrastructure (AMI) data. Most utility AMI measurements are recorded at hourly intervals, with only a small portion at higher resolutions, and even fewer households have appliance-level energy usage data. To address this issue, this paper proposes a two-stage long short-term memory (LSTM) framework for estimating household flexibility envelopes from low-resolution AMI data. In the first stage, the heating, ventilating, and air-conditioning (HVAC) load and non-HVAC loads are estimated by using a model trained on a small set of households with appliance-level profiles. These estimated data are then used to compute the upper- and lower-flexibility bounds, which are subsequently down-sampled to lower-resolution data. In the second stage, these flexibility bounds serve as training inputs for another LSTM model, enabling direct prediction of flexibility envelopes for households with only hourly AMI data. This method is validated using Pecan Street data from two different areas, and the results demonstrate its applicability and effectiveness.

24 POWER TRANSMISSION AND DISTRIBUTION

Cloud-Based Orchestration of a Model-Based Power and Data Analysis Toolchain

The proposed Europa Mission concept contains many engineering and scientific instruments that consume varying amounts of power and produce varying amounts of data throughout the mission. System-level power and data usage must be well understood and analyzed to verify design requirements. Numerous cross-disciplinary tools and analysis models are used to simulate the system-level spacecraft power and data behavior. This paper addresses the problem of orchestrating a consistent set of models, tools, and data in a unified analysis toolchain when ownership is distributed among numerous domain experts. An analysis and simulation environment was developed as a way to manage the complexity of the power and data analysis toolchain and to reduce the simulation turnaround time. A system model data repository is used as the trusted store of high-level inputs and results while other remote servers are used for archival of larger data sets and for analysis tool execution. Simulation data passes through numerous domain-specific analysis tools and end-to-end simulation execution is enabled through a web-based tool. The use of a cloud-based service facilitates coordination among distributed developers and enables scalable computation and storage needs, and ensures a consistent execution environment. Configuration management is emphasized to maintain traceability between current and historical simulation runs and their corresponding versions of models, tools and data.

Post, Ethan

Catalyzing deep decarbonization with federated battery diagnosis and prognosis for better data management in energy storage systems

Industrial data analytics methods play a central role in improving energy storage performance and efficiency, impacting the future of electrified transportation and renewable electricity generation. However, significant challenges hinder the large-scale deployment of batteries. Conventional methods rely on centralized collection and processing of fleet-level data, leading to database size issues and privacy concerns due to potential data breaches. To enable scalable deployment of battery management systems, this article proposes a federated battery diagnosis and prognosis model, which distributes the processing of battery standard current-voltage-time-usage data in a privacy-preserving manner. Instead of transferring the raw data, this approach communicates only the locally processed parameters, thus reducing communication load and preserving data confidentiality. The federated model offers a paradigm shift in battery health management through privacy-preserving distributed methods for battery data processing and lifetime prediction, ensuring the reliable and sustainable deployment of lithium-ion batteries in a rapidly evolving world.

asset health management