Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Large Dataset Processing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Data-driven picosecond X-ray imaging for quantitative plasma-induced shock characterization

Imaging dynamic events, especially shockwave behavior, is key to advancing high-energy-density (HED) research. Recent advances in fourth- and fifth-generation X-ray light sources allow for high-resolution imaging of fast phenomena, but limited beam time necessitates maximizing data acquisition. We present a benchtop-scale pulsed plasma device submerged in liquid heptane, capable of generating dynamic events at rates exceeding 10 Hz, supporting the field’s data-driven goals by producing large, high-quality imaging datasets. Using X-ray phase contrast imaging (XPCI) at the Advanced Photon Source, we imaged weak shockwaves (Mach ~ 1.2) in heptane interacting with plasma-induced cavitation bubbles, causing deviation from Rankine-Hugoniot behavior; to our knowledge, this represents the first direct imaging of such interaction. Our quantitative analysis offers insight into weak shock phenomena and energy-focusing applications in pulsed plasmas. These results highlight the potential for large datasets to advance dynamic HED research at current light source facilities, and have implications for fields such as inertial confinement fusion, plasma-enhanced chemical processing, and biomedical applications.

36 MATERIALS SCIENCE↗

Towards an Aviation Large Language Model by Fine-tuning and Evaluating Transformers

In the aviation domain, there are many applications for machine learning and artificial intelligence tools that utilize natural language. For example, there is a desire to know the commonalities in written safety reports such as voluntary post incidents reports or create more accurate transcripts of air traffic management conversations. Another use-case is the possibility of extracting airspace procedures and constraints currently written in documents such as Letters of Agreement (LOA) which is used as the evaluation case in this paper. These applications can benefit from the use of state-of-the-art Natural Language Processing (NLP) techniques when adapted to the language/phraseology specific to the aviation domain. This paper evaluates the viability of transferring pre-trained large language models to the aviation domain by adapting transformer based models using aviation datasets. This paper utilized two datasets to adapt a ‘Robustly Optimized Bidirectional Encoder Representations from Transformers Approach’ (RoBERTa) model and two down-stream classification tasks to assess its performance. These datasets are all built upon Letters of Agreement which are Federal Aviation Administration (FAA) documents that formalize airspace operations across the national airspace system. The first two datasets are used for the adaptation of RoBERTa to the aviation domain and were of different sizes to assess the number of documents needed to adapt to the aviation domain. They contain many examples of ‘aviation English’ using domain specific terminology and phrasing which serves as a representative basis to perform the unsupervised adaptation. The second dataset is a separate set of LOA documents with two sets of classification labels to be used for evaluation; one at the document level and one at the line level. These down-stream evaluations allowed the measurement of improvement by adapting RoBERTa. The accuracy increased by 4-6% on both tasks and the F1 score on the class of interest increased by 4-8% from the adaptation.

Air Traffic Management↗

Towards an Aviation Large Language Model by Fine-tuning and Evaluating Transformers

In the aviation domain, there are many applications for machine learning and artificial intelligence tools that utilize natural language. For example, there is a desire to know the commonalities in written safety reports such as voluntary post incidents reports or create more accurate transcripts of air traffic management conversations. Another use-case is the possibility of extracting airspace procedures and constraints currently written in documents such as Letters of Agreement (LOA) which is used as the evaluation case in this paper. These applications can benefit from the use of state-of-the-art Natural Language Processing (NLP) techniques when adapted to the language/phraseology specific to the aviation domain. This paper evaluates the viability of transferring pre-trained large language models to the aviation domain by adapting transformer based models using aviation datasets. This paper utilized two datasets to adapt a ‘Robustly Optimized Bidirectional Encoder Representations from Transformers Approach’ (RoBERTa) model and two down-stream classification tasks to assess its performance. These datasets are all built upon Letters of Agreement which are Federal Aviation Administration (FAA) documents that formalize airspace operations across the national airspace system. The first two datasets are used for the adaptation of RoBERTa to the aviation domain and were of different sizes to assess the number of documents needed to adapt to the aviation domain. They contain many examples of ‘aviation English’ using domain specific terminology and phrasing which serves as a representative basis to perform the unsupervised adaptation. The second dataset is a separate set of LOA documents with two sets of classification labels to be used for evaluation; one at the document level and one at the line level. These down-stream evaluations allowed the measurement of improvement by adapting RoBERTa. The accuracy increased by 4-6% on both tasks and the F1 score on the class of interest increased by 4-8% from the adaptation.

Air Traffic Management↗

Processing of DMSP magnetic data: Handbook of programs, tapes, and datasets

The DMSP F-7 satellite was an operational Air Force meteorological satellite which carried a magnetometer for geophysical measurements. The magnetometer was located within the body of the spacecraft in the presence of large spacecraft fields. In addition to stray magnetic fields, the data have inherent position and time inaccuracies. Algorithms were developed to identify and remove time varying magnetic field noise from the data. These algorithms are embodied in an automated procedure which fits a smooth curve through the data and then identifies outliers and which filters the predominant Fourier component of noise from the data. Techniques developed for Magsat were then modified and used to attempt determination of the spacecraft fields, of any rotation between the magnetometer axes and the spacecraft axes, and of any scale changes within the magnetometer itself. Software setup and usage are documented and program listings are included in the Appendix. The initial and resulting data are archived on magnetic cartridge and the formats are documented.

Langel, R. A.↗

DECOVALEX-2023: Task C Final Report

The Full-scale Emplacement (FE) heater experiment at the Mont Terri Underground Rock Laboratory (URL) was designed and conducted by Nagra to replicate an emplacement tunnel of Nagra’s reference repository design at 1:1 scale. Alongside testing the technical feasibility of constructing disposal tunnels, emplacing waste containers in the tunnels and then backfilling them, the main goals of the FE experiment are (1) to obtain a better understanding of the coupled effects of induced thermo-hydro-mechanical (THM) processes that may occur and (2) to validate existing coupled THM models (Müller et al., 2017). A key aspect of ensuring safety for repositories located in low-permeability rock involves minimizing any damage to the rock itself, thereby preserving its integrity and promoting a stable environment Amongst a number of processes that could damage the rock is the increase in pore pressure due to thermal loading caused by heat emitted from the waste. To reduce the potential damage of the rock, it is important to analyse the evolution of heat over time due to the heat load of the containers and assess possible consequences by coupled THM models. The aim of Task C of DECOVALEX-2023 was to build 3D numerical models of the FE experiment, focussing in particular on the heating induced pore pressure change in the Opalinus Clay. Data from a large number of sensors were available from the FE experiment for model comparison. These sensors measured temperature and relative humidity in the bentonite around the heaters, and temperature, pressure and displacement/strain in the surrounding Opalinus clay. Data were available from the start of excavation (April 2012) up to August 2020 for most sensors (more than 5 years from the start of heating in December 2014). To fulfil the overall aim of the task, the work was broken down into a number of steps, starting with simpler models to build confidence in each team’s approach and then moving to more complex models that better represent the FE experiment. Step 0 consisted of 2D benchmark models, gradually increasing the number of processes that are represented from thermal (T) only models in Step 0a, to coupled thermal hydraulic (TH) models in Step 0b with a representation of changing porosity, to coupled thermo-hydro-mechanical (THM) models in Step 0c, where porosity changes are calculated by the mechanical model. A detailed specification of processes, parameters, initial and boundary conditions was provided for this step, with the ambition that all teams would work towards close agreement in their model results, thus building confidence in the model implementations. vi It was not straightforward to achieve agreement between the teams, so additional steps (Step 0b2, 0b3, 0c2, 0c3) were added along with derivation of some analytical solutions against which the models could be compared. The reasons for the differences between teams were investigated and found to be caused primarily by different conceptual model assumptions (including temperature dependence of the thermal expansion of water), different model formulations (including porosity evolution) and differences in modelled domain sizes, boundary conditions and grid discretisation. This demonstrates that comparisons between multiple modelling teams and/or comparison with analytical results and experimental data are highly beneficial in providing an indication of uncertainty in model predictions. At the conclusion of Step 0, almost all teams had achieved a close agreement in model results and those that had not achieved an agreement knew the reason for this. Step 1 moved from 2D models to 3D models of the FE experiment without adding technical features like shotcrete or EDZ, and only considering the heating phase. Initially the 3D model was tightly specified to continue to build confidence in the model implementations (Step 1a). The results of Step 1a were compared to the data from the FE-experiment without the teams seeing the data. The teams were then provided with a sub-set of the data from the FE-experiment and invited to consider how best to use the large dataset for model comparison (Step 1b). Teams were then asked to use the data provided to calibrate their models, only changing material property values rather than adding features or processes to their models (Step 1c). In Step 1, teams were asked to only model the heating phase of the experiment, so pressure in the Opalinus Clay was reported as change in pressure since the initial conditions were specified rather than modelled. The change from 2D to 3D models was accompanied by an increase in the dispersion of results between the teams. Some of this was resolved during the task, but some remained and is potentially due to model discretisation. Calibration of parameters was useful in improving the fit of the models to the data but the remaining differences indicated that the models were missing features or processes. In Step 2, the teams were asked to update their models with additional features and processes as well as calibrating parameters to try and improve the fit of the models to the data. Teams were encouraged to represent ventilation of the open FE tunnel prior to backfilling with heaters and bentonite and in Step 2, the absolute pressure in the Opalinus Clay was compared between the teams. Teams took different approaches, but there was consideration of adding shotcrete and an EDZ into the model, representing stress change during excavation and different approaches to modelling ventilation of the FE tunnel. Overall, the documented results showed a very good agreement for temperature. The results for porewater pressure evolution showed a significant improvement for most teams compared to Step 1c with a good agreement to the measurements for several teams whereas some teams overpredicted the pressure increase and others overpredicted the drainage effect especially for the sensors close to the heater. Step 3 was an opportunity for teams to use the models developed in Step 1 and Step 2 to make predictions about the temperature and pressure changes that will be expected at the FE experiment over the next few years in light of the planned changes in thermal output of the heaters.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

Simultaneous lidar profiling of water vapor, aerosol, and wind over the eastern Pacific and implications for observational challenges on clouds, circulation, and climate sensitivity

Moist convective processes and related circulations are key components of the Earth’s climate system. Improving our understanding of moisture and cloud processes over oceans is especially challenging due to observational limitations, but also important due to the large radiative effects of cloud over ocean as well as the ties to local and global-scale circulations. A unique dataset for addressing these topics was collected over 5 flights with the NASA HALO (water vapor DIAL and aerosol HSRL) and DAWN (Doppler wind) lidar systems onboard the NASA DC-8 aircraft during the spring 2019 ADM-Aeolus Cal/Val test flight campaign. High-resolution water vapor, aerosol/cloud, and wind profiles were simultaneously collected and complimented with in situ high-resolution dropsonde measurements of temperature, relative humidity, and wind profiles. Airborne lidar measurements of these key variables captured both large-scale spatial variability across different cloud and atmospheric conditions as well as finer features, providing insight into processes from large-scale circulations down to aerosol/cloud interactions. The 5 flights spanned the northeast Pacific and southwest U.S. from 6 N to 52 N and 112 W to 157 W and captured several features of interest including moist layers from the ascending region of the Hadley circulation, very dry layers above midlatitude marine stratocumulus and cloud-free planetary boundary layers, and cloud top height and spatial variability. High covariance of aerosol and water vapor in the free troposphere was also observed for the first time. Preliminary analysis of these data sets will be presented as they apply to clouds and circulations, supplemented by model reanalysis to inform how regional water vapor structure and large-scale motion may contribute to observed shallow cloud formations and other meteorological processes that underpin our current understanding of how these state variables couple to impact the Earth’s weather and climate systems. The potential of complementary measurements in future field campaigns will also be mentioned.

Brian James Carroll↗

Digital Lunar Exploration Sites (DLES) Terrain Crafting

Humans will soon be returning to the surface of the Moon with NASA’s Artemis program. The Artemis program is an international collaboration that will consist of a complex series of space systems and missions to explore the lunar surface and pave the way for the future exploration of Mars. NASA and its partners rely heavily on simulation for lighting and navigation studies as well as training astronauts, flight controllers, and mission support staff. The NASA Exploration Systems Simulations (NExSyS) team in the Simulation and Graphics Branch (ER7) in the Engineering Directorate at NASA’s Johnson Space Center has built up many simulation products to support this effort, one of which is the Digital Lunar Exploration Sites (DLES). DLES is a collection of products used to simulate and render the lunar surface in a digital environment. We discussed and presented an overview of the DLES products at the 2022 IEEE Aerospace Conference in Big Sky, MT with a paper titled "Digital Lunar Exploration Sites". This “DLES Terrain Crafting” paper will expand on the information previously provided in “DLES” paper and dive deeper into the details of the terrain crafting process and the toolsets used to support this task. The best digital data currently available of the lunar surface is provided by the Lunar Reconnaissance Orbiter (LRO). Its Lunar Orbiter Laser Altimeter (LOLA) achieves an impressive resolution of 5m per pixel at the Lunar South Pole (LSP) and can generate datasets covering a large continuous region near the LSP. There are a few additional methods, such as Shape from Shading which can infer higher resolution data (up to 1m per pixel) from the LRO Narrow Angle Camera (NAC) images. However, surface-based simulations require higher-resolution data, and this paper will discuss the process of enhancing the terrain to meet that need. The process begins with capturing statistical data of craters in the regions of interest using images provided by the LRO NAC. This data is then used to scatter artificial features which are not captured in the truth data, resulting in an enhanced DEM with a much higher resolution of 20cm per pixel. Many tools were built up to assist in the creation of these artificial Digital Elevation Models (DEM), which this paper will discuss in detail. DEMs themselves are a very powerful representation of a planetary surface, and many operations and tools can utilize the data they contain. This paper includes a description of the rendering of the lunar surface in a graphics engine, generation of contact patches to simulate tire to ground interaction, and ray tracing utilities to model Line of Sight (LOS) interactions with the terrain. This paper will also explore some new tool sets currently under development which aim to utilize Machine Learning (ML) to assist in the identification of craters from LRO NAC imagery. While this is not a novel idea, the NExSyS team is developing a unique approach which may result in more robust identification of crater characteristics.

Artemis↗

Spatiotemporal Downscaling Model for Solar Irradiance Forecast Using Nearest-Neighbor Random Forest and Gaussian Process

Accurate solar photovoltaic (PV) capacity estimation requires high-resolution, site-specific solar irradiance data to account for localized variability. However, global datasets, such as the National Solar Radiation Database (NSRDB), provide regional averages that fail to capture the fine-scale fluctuations critical for large-scale grid integration. This limitation is particularly relevant in the context of increasing distributed energy resources (DERs) penetration, such as rooftop PV. Additionally, it is critical to the implementation of the U.S. Federal Energy Regulatory Commission (FERC) Order 2222, which facilitates DER participation in U.S. bulk power markets. To address this challenge, this study evaluates Nearest-Neighbor Random Forest (NNRF) and Nearest-Neighbor Gaussian Process (NNGP) models for spatiotemporal downscaling of global solar irradiance data. By leveraging historical irradiance and meteorological data, these models incorporate spatial, temporal, and feature-based correlations to enhance local irradiance predictions. The NNRF model, a machine-learning approach, prioritizes computational efficiency and predictive accuracy, while the NNGP model offers a level of interpretability and prediction uncertainty by numerically quantifying correlations and dependencies in the data. Model validation was conducted using day-ahead predictions. The results showed that the average Goodness of Fit (GoF) of the NNRF model of 90.61% across all eight sites outperformed the GoF of the NNGP of 85.88%. Additionally, the computational speed of NNRF was 2.5 times faster than the NNGP. Finally, the NNGP displayed polynomial scaling while the NNRF scaled linearly with increasing number of nearest neighbors. Additional validation of the model on five sites in Puerto Rico further confirmed the superiority of the NNRF model over the NNGP model. These findings highlight the robustness and computational efficiency of NNRF for large-scale solar irradiance downscaling, making it a strong candidate for improving PV capacity estimation and real-time electricity market integration for DERs.

Asiedu, Shadrack (ORCID:0009000646004826)↗

Automated Knowledge Discovery From Simulators

A computational method, SimLearn, has been devised to facilitate efficient knowledge discovery from simulators. Simulators are complex computer programs used in science and engineering to model diverse phenomena such as fluid flow, gravitational interactions, coupled mechanical systems, and nuclear, chemical, and biological processes. SimLearn uses active-learning techniques to efficiently address the "landscape characterization problem." In particular, SimLearn tries to determine which regions in "input space" lead to a given output from the simulator, where "input space" refers to an abstraction of all the variables going into the simulator, e.g., initial conditions, parameters, and interaction equations. Landscape characterization can be viewed as an attempt to invert the forward mapping of the simulator and recover the inputs that produce a particular output. Given that a single simulation run can take days or weeks to complete even on a large computing cluster, SimLearn attempts to reduce costs by reducing the number of simulations needed to effect discoveries. Unlike conventional data-mining methods that are applied to static predefined datasets, SimLearn involves an iterative process in which a most informative dataset is constructed dynamically by using the simulator as an oracle. On each iteration, the algorithm models the knowledge it has gained through previous simulation trials and then chooses which simulation trials to run next. Running these trials through the simulator produces new data in the form of input-output pairs. The overall process is embodied in an algorithm that combines support vector machines (SVMs) with active learning. SVMs use learning from examples (the examples are the input-output pairs generated by running the simulator) and a principle called maximum margin to derive predictors that generalize well to new inputs. In SimLearn, the SVM plays the role of modeling the knowledge that has been gained through previous simulation trials. Active learning is used to determine which new input points would be most informative if their output were known. The selected input points are run through the simulator to generate new information that can be used to refine the SVM. The process is then repeated. SimLearn carefully balances exploration (semi-randomly searching around the input space) versus exploitation (using the current state of knowledge to conduct a tightly focused search). During each iteration, SimLearn uses not one, but an ensemble of SVMs. Each SVM in the ensemble is characterized by different hyper-parameters that control various aspects of the learned predictor - for example, whether the predictor is constrained to be very smooth (nearby points in input space lead to similar output predictions) or whether the predictor is allowed to be "bumpy." The various SVMs will have different preferences about which input points they would like to run through the simulator next. SimLearn includes a formal mechanism for balancing the ensemble SVM preferences so that a single choice can be made for the next set of trials.

Burl, Michael↗

The NAS Computational Aerosciences Archive

In order to further the state-of-the-art in computational aerosciences (CAS) technology, researchers must be able to gather and understand existing work in the field. One aspect of this information gathering is studying published work available in scientific journals and conference proceedings. However, current scientific publications are very limited in the type and amount of information that they can disseminate. Information is typically restricted to text, a few images, and a bibliography list. Additional information that might be useful to the researcher, such as additional visual results, referenced papers, and datasets, are not available. New forms of electronic publication, such as the World Wide Web (WWW), limit publication size only by available disk space and data transmission bandwidth, both of which are improving rapidly. The Numerical Aerodynamic Simulation (NAS) Systems Division at NASA Ames Research Center is in the process of creating an archive of CAS information on the WWW. This archive will be based on the large amount of information produced by researchers associated with the NAS facility. The archive will contain technical summaries and reports of research performed on NAS supercomputers, visual results (images, animations, visualization system scripts), datasets, and any other supporting meta-information. This information will be available via the WWW through the NAS homepage, located at http://www.nas.nasa.gov/, fully indexed for searching. The main components of the archive are technical summaries and reports, visual results, and datasets. Technical summaries are gathered every year by researchers who have been allotted resources on NAS supercomputers. These summaries, together with supporting visual results and references, are browsable by interested researchers. Referenced papers made available by researchers can be accessed through hypertext links. Technical reports are in-depth accounts of tools and applications research projects performed by NAS staff members and collaborators. Visual results, which may be available in the form of images, animations, and/or visualization scripts, are generated by researchers with respect to a certain research project, depicting dataset features that were determined important by the investigating researcher. For example, script files for visualization systems (e.g. FAST, PLOT3D, AVS) are provided to create visualizations on the user's local workstation to elucidate the key points of the numerical study. Users can then interact with the data starting where the investigator left off. Datasets are intended to give researchers an opportunity to understand previous work, 'mine' solutions for new information (for example, have you ever read a paper thinking "I wonder what the helicity density looks like?"), compare new techniques with older results, collaborate with remote colleagues, and perform validation. Supporting meta-information associated with the research projects is also important to provide additional context for research projects. This may include information such as the software used in the simulation (e.g. grid generators, flow solvers, visualization). In addition to serving the CAS research community, the information archive will also be helpful to students, visualization system developers and researchers, and management. Students (of any age) can use the data to study fluid dynamics, compare results from different flow solvers, learn about meshing techniques, etc., leading to better informed individuals. For these users it is particularly important that visualization be integrated into dataset archives. Visualization researchers can use dataset archives to test algorithms and techniques, leading to better visualization systems, Management can use the data to figure what is really going on behind the viewgraphs. All users will benefit from fast, easy, and convenient access to CFD datasets. The CAS information archive hopes to serve as a useful resource to those interested in computational sciences. At present, only information that may be distributed internationally is made available via the archive. Studies are underway to determine security requirements and solutions to make additional information available. By providing access to the archive via the WWW, the process of information gathering can be more productive and fruitful due to ease of access and ability to manage many different types of information. As the archive grows, additional resources from outside NAS will be added, providing a dynamic source of research results.

Miceli, Kristina D.↗

Visualizing UPSP Data with Python

The Unsteady Pressure-Sensitive Paint (uPSP) projects uses Pressure-Sensitive paint applied over aerospace models during wind tunnel testing to collect pressure data with high spatial and temporal resolution in order to inform unsteady aerodynamics studies. For each of the 800+ experimental runs, four cameras generate up to 50 GB of video data, which must then be processed, analyzed, and visualized on the NASA Advanced Supercomputing system (NAS) to assess the result. One of the final data analysis products is the dynamic modal decomposition (DMD) results, which decomposes the pressure reading signals by their frequency component. The goal of this project is to visualize the DMD results over a 3D rendering of the model, using efficient and parallelized python routines. The software uses the pytecplot library, a high-level API that connects python scripting to a Tecplot 360 engine. Tecplot is an industry standard high-performance visualization tool that can handle large datasets and workflow. Various animation, rendering, and image-combination techniques were investigated to generate the final videos using OpenCV on the NAS. The final result is a software tool that takes in data products from the uPSP processing chain and generates high resolution visualization videos in parallel for every data file, allowing researchers to view their results efficiently and at an unprecedentedly detailed level.

Emma Dolores McMillian↗

Learning together: Towards foundation models for machine learning interatomic potentials with meta-learning

Abstract The development of machine learning models has led to an abundance of datasets containing quantum mechanical (QM) calculations for molecular and material systems. However, traditional training methods for machine learning models are unable to leverage the plethora of data available as they require that each dataset be generated using the same QM method. Taking machine learning interatomic potentials (MLIPs) as an example, we show that meta-learning techniques, a recent advancement from the machine learning community, can be used to fit multiple levels of QM theory in the same training process. Meta-learning changes the training procedure to learn a representation that can be easily re-trained to new tasks with small amounts of data. We then demonstrate that meta-learning enables simultaneously training to multiple large organic molecule datasets. As a proof of concept, we examine the performance of a MLIP refit to a small drug-like molecule and show that pre-training potentials to multiple levels of theory with meta-learning improves performance. This difference in performance can be seen both in the reduced error and in the improved smoothness of the potential energy surface produced. We therefore show that meta-learning can utilize existing datasets with inconsistent QM levels of theory to produce models that are better at specializing to new datasets. This opens new routes for creating pre-trained, foundation models for interatomic potentials.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Multi-Stage System for Automatic Target Recognition

A multi-stage automated target recognition (ATR) system has been designed to perform computer vision tasks with adequate proficiency in mimicking human vision. The system is able to detect, identify, and track targets of interest. Potential regions of interest (ROIs) are first identified by the detection stage using an Optimum Trade-off Maximum Average Correlation Height (OT-MACH) filter combined with a wavelet transform. False positives are then eliminated by the verification stage using feature extraction methods in conjunction with neural networks. Feature extraction transforms the ROIs using filtering and binning algorithms to create feature vectors. A feedforward back-propagation neural network (NN) is then trained to classify each feature vector and to remove false positives. The system parameter optimizations process has been developed to adapt to various targets and datasets. The objective was to design an efficient computer vision system that can learn to detect multiple targets in large images with unknown backgrounds. Because the target size is small relative to the image size in this problem, there are many regions of the image that could potentially contain the target. A cursory analysis of every region can be computationally efficient, but may yield too many false positives. On the other hand, a detailed analysis of every region can yield better results, but may be computationally inefficient. The multi-stage ATR system was designed to achieve an optimal balance between accuracy and computational efficiency by incorporating both models. The detection stage first identifies potential ROIs where the target may be present by performing a fast Fourier domain OT-MACH filter-based correlation. Because threshold for this stage is chosen with the goal of detecting all true positives, a number of false positives are also detected as ROIs. The verification stage then transforms the regions of interest into feature space, and eliminates false positives using an artificial neural network classifier. The multi-stage system allows tuning the detection sensitivity and the identification specificity individually in each stage. It is easier to achieve optimized ATR operation based on its specific goal. The test results show that the system was successful in substantially reducing the false positive rate when tested on a sonar and video image datasets.

Chao, Tien-Hsin↗

Bridging the Last Mile with Open-Source Advancements: Empowering Communities through Fusion of Aerosol Optical Depth (AOD) Products from Multi-Satellite Sensors

Aerosol Optical Depth (AOD) is a crucial parameter for understanding atmospheric aerosol distribution and their impact on climate and air quality. With the growing number of Earth observation satellites, there is an abundance of AOD products derived from various sensors onboard both geostationary and low-orbit satellites. The availability of multiple datasets provides an opportunity to harness the strengths of each sensor and create comprehensive and accurate AOD datasets for climate and air quality studies at different temporal and spatial scales. Our NASA aerosol MEaSURES project has made significant strides in recent years by undertaking the ambitious task of developing an open-source package tailored for fusing AOD products from different sources. The package is based on OOP (Object-Oriented Programming) design and is implemented in Python modules. Generic interfaces enable easy inclusion of large and heterogeneous data. The package may be utilized to produce harmonized AOD datasets with enhanced spatial and temporal coverage. The latest version of the package is able to process and integrate the dark-target AOD data from six different sensors: AHI Himawari-8, ABI GOES-West, ABI GOES-East, MODIS AQUA, MODIS TERRA, and VIIRS SNPP. Rigorous validation and intercomparison studies have been performed to assess the accuracy and reliability of the fused AOD product against ground-based measurements and reference datasets. The open-source nature of the developed package ensures transparency, reproducibility, and community engagement. The research community and stakeholders can access, contribute to, and further improve the fusion methodology, making it adaptable to other studies, or expanding it to include new satellite data as they become available. In this poster presentation, we will introduce the accomplishments and challenges faced during the development of the open-source package for AOD data fusion, and demonstrate the advantages of combining AOD products from the six aforementioned satellite sensors. The presentation aims to foster discussions, collaborations, and future directions in integrating Earth observation and remote sensing data, which may contribute to a better understanding of atmospheric aerosols and their impacts on our environment.

Zhaohui Zhang↗

Access NASA Satellite Global Precipitation Data Visualization on YouTube

Since the satellite era began, NASA has collected a large volume of Earth science observations for research and applications around the world. The collected and archived satellite data at 12 NASA data centers can also be used for STEM education and activities such as disaster events, climate change, etc. However, accessing satellite data can be a daunting task for non-professional users such as teachers and students because of unfamiliarity of terminology, disciplines, data formats, data structures, computing resources, processing software, programming languages, etc. Over the years, many efforts including tools, training classes, and tutorials have been developed to improve satellite data access for users, but barriers still exist for non-professionals. In this presentation, we will present our latest activity that uses a very popular online video sharing Web site, YouTube (https://www.youtube.com/), for accessing visualizations of our global precipitation datasets at the NASA Goddard Earth Sciences (GES) Data and Information Services Center (DISC). With YouTube, users can access and visualize a large volume of satellite data without the necessity to learn new software or download data. The dataset in this activity is a one-month animation for the GPM (Global Precipitation Measurement) Integrated Multi-satellite Retrievals for GPM (IMERG). IMERG provides precipitation on a near-global (60 deg. N-S) coverage at half-hourly time interval, providing more details on precipitation processes and development compared to the 3-hourly TRMM (Tropical Rainfall Measuring Mission) Multisatellite Precipitation Analysis (TMPA, 3B42) product. When the retro-processing of IMERG during the TRMM era is finished in 2018, the entire video will contain more than 330,000 files and will last ~3.6 hours. Future plans include development of flyover videos for orbital data for an entire satellite mission or project. All videos, including the one-month animation, will be uploaded and available at the GES DISC site on YouTube (https://www.youtube.com/user/NASAGESDISC).

precipitation↗

Discovering the Unknowns: A First Step

This article aims at discovering the unknown variables in the system through data analysis. The main idea is to use the time of data collection as a surrogate variable and try to identify the unknown variables by modeling gradual and sudden changes in the data. We use Gaussian process modeling and a sparse representation of the sudden changes to efficiently estimate the large number of parameters in the proposed statistical model. The method is tested on a realistic dataset generated using a one-dimensional implementation of a Magnetized Liner Inertial Fusion (MagLIF) simulation model, and encouraging results are obtained.

42 ENGINEERING↗

Fostering Geothermal Machine Learning Success: Elevating Big Data Accessibility and Automated Data Standardization in the Geothermal Data Repository

The Department of Energy's (DOE) Geothermal Data Repository (GDR) has implemented improvements to both its data lakes and its data standards and automated data pipelines. The GDR data lakes have reduced storage and compute-related barriers to using large geothermal datasets, enabling these large datasets to be accessed by anyone with a modern computer and internet access. More recently, the GDR has been working to further reduce barriers through streamlining the data intake process, educating users on the process and requirements, and aiding users in accessing data from the data lakes. These improvements have augmented the quantity of datasets the GDR is able to accept into its data lakes and have enabled users who are new to cloud tools to access these datasets more easily, overall increasing the accessibility of big geothermal data for use in machine learning and other projects. In addition, the GDR now has built-in data standards and pipelines for drilling data, geospatial data, and distributed acoustic sensing (DAS) data. These standardization efforts aim to enhance the real-world applicability of geothermal machine learning outcomes by improving the quality of training data. Specifically, through standardizing high-value datasets, the GDR is reducing project-specific data curation requirements, thus allowing more time for actual research. By automating this process, the burden of standardization is lifted from the user, ultimately increasing the availability of standardized data.

15 GEOTHERMAL ENERGY↗

Fostering Geothermal Machine Learning Success: Elevating Big Data Accessibility and Automated Data Standardization in the Geothermal Data Repository: Preprint

The Department of Energy's (DOE) Geothermal Data Repository (GDR) has implemented improvements to both its data lakes and its data standards and automated data pipelines. The GDR data lakes have reduced storage and compute-related barriers to using large geothermal datasets, enabling these large datasets to be accessed by anyone with a modern computer and internet access. More recently, the GDR has been working to further reduce barriers through streamlining the data intake process, educating users on the process and requirements, and aiding users in accessing data from the data lakes. These improvements have augmented the quantity of datasets the GDR is able to accept into its data lakes and have enabled users who are new to cloud tools to access these datasets more easily, overall increasing the accessibility of big geothermal data for use in machine learning and other projects. In addition, the GDR now has built-in data standards and pipelines for drilling data, geospatial data, and distributed acoustic sensing (DAS) data. These standardization efforts aim to enhance the real-world applicability of geothermal machine learning outcomes by improving the quality of training data. Specifically, through standardizing high-value datasets, the GDR is reducing project-specific data curation requirements, thus allowing more time for actual research. By automating this process, the burden of standardization is lifted from the user, ultimately increasing the availability of standardized data.

accessibility↗