Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “large data sets”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Cosmochemical Studies: Meteorites, Asteroidal Processes, Chondrules

Our research mainly concerned the asteroidal processes involved in the formation of meteorites and meteoritic chondrules. We continued to generate large amounts of instrumental-neutron-activation analysis (INAA) data, both for irons, chondrites and primitive achondrites. Major themes of our chondrule research were: (1) the temperature and crystallization history of individual chondrules, and (2) the evolution of the solar nebula during the period within which chondrule formation occurred. Much of our chondrule research was focused on the highly primitive CO3 chondrites. We initiated a study of the cooling history of high-FeO chondrules by characterizing the overgrowth layers on relict grains.. We also continued our studies of the composition and the formation of iron meteorites and the evolution of their parent planets. The large data sets that we have generated at UCLA allows systematic comparisons of the large magmatic groups both in terms of fractional crystallization (including rough estimates of non-metal contents of the parental melts) and in terms of the effects of variable contents of trapped melt. We have completed a preliminary study of group IIIAB in which we developed a trapped-melt model and more detailed studies of group IVA and the main-group pallasites. By comparing these large groups and modeling them by a combination of crystallization and melt trapping, we are able to better define both the formation processes and the nature of the solid/liquid elemental partitioning. We helped maintain the excellent neutron-activation facilities at UCLA, a major resource for the cosmochemical community.

Wasson, John T.↗

Machine Learning Prediction of Tritium‐Helium Groundwater Ages in the Central Valley, California, USA

Abstract Groundwater ages provides insight into recharge rates, flow velocities, and vulnerability to contaminants. The ability to predict groundwater ages based on more accessible parameters via Machine Learning (ML) would advance our ability to guide sustainable management of groundwater resources. In this study, ML models were trained and tested on a large data set of tritium concentrations and tritium‐helium groundwater ages from the California Central Valley, a large groundwater basin with complex land use, irrigation, and water management practices. The ML models were trained on 63 features, including location, well construction information, landscape characteristics, and climate variables, water chemistry, and stable isotopes. The Bagging regressor method can accurately classify (F1‐score = 0.91) groundwater samples as either modern or pre‐modern whereas the accuracy of the ML prediction of continuous tritium‐helium groundwater ages is limited and explains only of the variability in this data set. In general, ML groundwater age prediction relies mostly on features related to (a) the source of groundwater recharge, (b) contaminant history, (c) aquifer materials, (d) well construction, and (e) geochemical reactions along flow paths.

54 ENVIRONMENTAL SCIENCES↗

Geographic information system for fusion and analysis of high-resolution remote sensing and ground truth data

We seek to combine high-resolution remotely sensed data with models and ground truth measurements, in the context of a Geographical Information System, integrated with specialized image processing software. We will use this integrated system to analyze the data from two Case Studies, one at a bore Al forest site, the other a tropical forest site. We will assess the information content of the different components of the data, determine the optimum data combinations to study biogeophysical changes in the forest, assess the best way to visualize the results, and validate the models for the forest response to different radar wavelengths/polarizations. During the 1990's, unprecedented amounts of high-resolution images from space of the Earth's surface will become available to the applications scientist from the LANDSAT/TM series, European and Japanese ERS-1 satellites, RADARSAT and SIR-C missions. When the Earth Observation Systems (EOS) program is operational, the amount of data available for a particular site can only increase. The interdisciplinary scientist, seeking to use data from various sensors to study his site of interest, may be faced with massive difficulties in manipulating such large data sets, assessing their information content, determining the optimum combinations of data to study a particular parameter, visualizing his results and validating his model of the surface. The techniques to deal with these problems are also needed to support the analysis of data from NASA's current program of Multi-sensor Airborne Campaigns, which will also generate large volumes of data. In the Case Studies outlined in this proposal, we will have somewhat unique data sets. For the Bonanza Creek Experimental Forest (Case I) calibrated DC-8 SAR data and extensive ground truth measurement are already at our disposal. The data set shows documented evidence to temporal change. The Belize Forest Experiment (Case II) will produce calibrated DC-8 SAR and AVIRIS data, together with extensive measurements on the tropical rain forest itself. The extreme range of these sites, one an Arctic forest, the other a tropical rain forest, has been deliberately chosen to find common problems which can lead to generalized observations and unique problems with data which raise issues for the EOS System.

Freeman, Anthony↗

Geographic information system for fusion and analysis of high-resolution remote sensing and ground data

We seek to combine high-resolution remotely sensed data with models and ground truth measurements, in the context of a Geographical Information System (GIS), integrated with specialized image processing software. We will use this integrated system to analyze the data from two Case Studies, one at a boreal forest site, the other a tropical forest site. We will assess the information content of the different components of the data, determine the optimum data combinations to study biogeophysical changes in the forest, assess the best way to visualize the results, and validate the models for the forest response to different radar wavelengths/polarizations. During the 1990's, unprecedented amounts of high-resolution images from space of the Earth's surface will become available to the applications scientist from the LANDSAT/TM series, European and Japanese ERS-1 satellites, RADARSAT and SIR-C missions. When the Earth Observation Systems (EOS) program is operational, the amount of data available for a particular site can only increase. The interdisciplinary scientist, seeking to use data from various sensors to study his site of interest, may be faced with massive difficulties in manipulating such large data sets, assessing their information content, determining the optimum combinations of data to study a particular parameter, visualizing his results and validating his model of the surface. The techniques to deal with these problems are also needed to support the analysis of data from NASA's current program of Multi-sensor Airborne Campaigns, which will also generate large volumes of data. In the Case Studies outlined in this proposal, we will have somewhat unique data sets. For the Bonanza Creek Experimental Forest (Case 1) calibrated DC-8 SAR (Synthetic Aperture Radar) data and extensive ground truth measurement are already at our disposal. The data set shows documented evidence to temporal change. The Belize Forest Experiment (Case 2) will produce calibrated DC-8 SAR and AVIRIS data, together with extensive measurements on the tropical rain forest itself. The extreme range of these sites, one an Arctic forest, the other a tropical rain forest, has been deliberately chosen to find common problems which can lead to generalized observations and unique problems with data which raise issues for the EOS System.

Freeman, Anthony↗

Deeply learning deep inelastic scattering kinematics

We study the use of deep learning techniques to reconstruct the kinematics of the neutral current deep inelastic scattering (DIS) process in electron–proton collisions. In particular, we use simulated data from the ZEUS experiment at the HERA accelerator facility, and train deep neural networks to reconstruct the kinematic variables Q 2 and x. Our approach is based on the information used in the classical construction methods, the measurements of the scattered lepton, and the hadronic final state in the detector, but is enhanced through correlations and patterns revealed with the simulated data sets. We show that, with the appropriate selection of a training set, the neural networks sufficiently surpass all classical reconstruction methods on most of the kinematic range considered. Rapid access to large samples of simulated data and the ability of neural networks to effectively extract information from large data sets, both suggest that deep learning techniques to reconstruct DIS kinematics can serve as a rigorous method to combine and outperform the classical reconstruction methods.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Physics Mining of Multi-Source Data Sets

Powerful new parallel data mining algorithms can produce diagnostic and prognostic numerical models and analyses from observational data. These techniques yield higher-resolution measures than ever before of environmental parameters by fusing synoptic imagery and time-series measurements. These techniques are general and relevant to observational data, including raster, vector, and scalar, and can be applied in all Earth- and environmental science domains. Because they can be highly automated and are parallel, they scale to large spatial domains and are well suited to change and gap detection. This makes it possible to analyze spatial and temporal gaps in information, and facilitates within-mission replanning to optimize the allocation of observational resources. The basis of the innovation is the extension of a recently developed set of algorithms packaged into MineTool to multi-variate time-series data. MineTool is unique in that it automates the various steps of the data mining process, thus making it amenable to autonomous analysis of large data sets. Unlike techniques such as Artificial Neural Nets, which yield a blackbox solution, MineTool's outcome is always an analytical model in parametric form that expresses the output in terms of the input variables. This has the advantage that the derived equation can then be used to gain insight into the physical relevance and relative importance of the parameters and coefficients in the model. This is referred to as physics-mining of data. The capabilities of MineTool are extended to include both supervised and unsupervised algorithms, handle multi-type data sets, and parallelize it.

Helly, John↗

Astronomaly at scale: searching for anomalies amongst 4 million galaxies

ABSTRACT Modern astronomical surveys are producing data sets of unprecedented size and richness, increasing the potential for high-impact scientific discovery. This possibility, coupled with the challenge of exploring a large number of sources, has led to the development of novel machine-learning-based anomaly detection approaches, such as astronomaly. For the first time, we test the scalability of astronomaly by applying it to almost 4 million images of galaxies from the Dark Energy Camera Legacy Survey. We use a trained deep learning algorithm to learn useful representations of the images and pass these to the anomaly detection algorithm isolation forest, coupled with astronomaly’s active learning method, to discover interesting sources. We find that data selection criteria have a significant impact on the trade-off between finding rare sources such as strong lenses and introducing artefacts into the data set. We demonstrate that active learning is required to identify the most interesting sources and reduce artefacts, while anomaly detection methods alone are insufficient. Using astronomaly, we find 1635 anomalies among the top 2000 sources in the data set after applying active learning, including eight strong gravitational lens candidates, 1609 galaxy merger candidates, and 18 previously unidentified sources exhibiting highly unusual morphology. Our results show that by leveraging the human–machine interface, astronomaly is able to rapidly identify sources of scientific interest even in large data sets.

Astronomy & Astrophysics↗

Color contouring for atmospheric data sets

A program has been developed for use by researchers at the Langley Research Center (LaRC) to allow a quick and easy method to display color contour plots. The bilinear interpolation technique used in the contouring routine enables the user to analyze large data sets with smoothly varying color changes instead of dashes and lines. Annotations can be added to enhance the contour plot, and the resulting plot can easily be made into a color print or a viewgraph.

Ferebee, M. T.↗

A variational encoder–decoder approach to precise spectroscopic age estimation for large Galactic surveys

Constraints on the formation and evolution of the Milky Way Galaxy require multidimensional measurements of kinematics, abundances, and ages for a large population of stars. Ages for luminous giants, which can be seen to large distances, are an essential component of studies of the Milky Way, but they are traditionally very difficult to estimate precisely for a large data set and often require careful analysis on a star-by-star basis in asteroseismology. Because spectra are easier to obtain for large samples, being able to determine precise ages from spectra allows for large age samples to be constructed, but spectroscopic ages are often imprecise and contaminated by abundance correlations. Here we present an application of a variational encoder–decoder on cross-domain astronomical data to solve these issues. The model is trained on pairs of observations from APOGEE and Kepler of the same star in order to reduce the dimensionality of the APOGEE spectra in a latent space while removing abundance information. The low dimensional latent representation of these spectra can then be trained to predict age with just ∼1000 precise seismic ages. We demonstrate that this model produces more precise spectroscopic ages (∼ 22 per cent overall, ∼ 11 per cent for red-clump stars) than previous data-driven spectroscopic ages while being less contaminated by abundance information (in particular, our ages do not depend on [α/M]). We create a public age catalogue for the APOGEE DR17 data set and use it to map the age distribution and the age-[Fe/H]-[α/M] distribution across the radial range of the Galactic disc.

79 ASTRONOMY AND ASTROPHYSICS↗

Contextual classification on the massively parallel processor

Classifiers are often used to produce land cover maps from multispectral Earth observation imagery. Conventionally, these classifiers have been designed to exploit the spectral information contained in the imagery. Very few classifiers exploit the spatial information content of the imagery, and the few that do rarely exploit spatial information content in conjunction with spectral and/or temporal information. A contextual classifier that exploits spatial and spectral information in combination through a general statistical approach was studied. Early test results obtained from an implementation of the classifier on a VAX-11/780 minicomputer were encouraging, but they are of limited meaning because they were produced from small data sets. An implementation of the contextual classifier is presented on the Massively Parallel Processor (MPP) at Goddard that for the first time makes feasible the testing of the classifier on large data sets.

Tilton, James C.↗

Extended testing of a general contextual classifier using the massively parallel processor - Preliminary results and test plans

Earlier encouraging test results of a contextual classifier that combines spatial and spectral information employing a general statistical approach are expanded. The earlier results were of limited meaning because they were produced from small (50-by-50 pixel) data sets. An implementation of the contextual classifier on NASA Goddard's Massively Parallel Processor (MPP) is presented; for the first time the MPP makes feasible the testing of the classifier on large data sets (a 12-hour test on a VAX-11/780 minicomputer now takes 5 minutes on the MPP). The MPP is a Single-Instruction, Multiple Data Stream computer, consisting of 16,384 bit serial microprocessors connected in a 128-by-128 mesh array with each element having data transfer connections with its four nearest neighbors so that the MPP is capable of billions of operations per second. Preliminary results are given (with more expected for the conference) and plans are mentioned for extended testing of the contextual classifier on Thematic Mapper data sets.

Tilton, J. C.↗

Potential origin of the state-dependent hard tail in the black hole microquasar Cygnus X-1 as seen with INTEGRAL

Context.0.1–10 MeV observations of the black hole microquasar Cygnus X-1 have shown the presence of a spectral feature in the form of a power law in addition to the standard black body (0.1–10 keV) and Comptontonization (10–200 keV) components usually seen in all black hole X-ray binaries. This so-called “high-energy tail” has recently been shown to be strong in the hard spectral state and has been interpreted as high-energy part of the emission from a compact jet. Aims. This result was, however, obtained from a data set largely dominated by hard state observations. In the soft state, only upper limits on the presence and hence the potential parameters of a hard tail could be derived. Using an extended data set we aim at obtaining better constraints on the properties of this spectral component in both states. Methods. We make use of data obtained from about 15 years of observations with the INTEGRAL satellite. The data set is separated into the different states and we analyse stacked state-resolved spectra obtained from both the gamma-ray Imager and the Spectrometer onboard. Results. A high-energy component is detected in both states, confirming its earlier detection in the hard state and its suspected presence in the soft state as seen in a much smaller SPI data set. We first characterize the hard tail components in the two states through a model-independent, phenomenological analysis. We then apply physical models based on hybrid Comptonization (eqpair and belm). The spectra are well modeled in all cases, with a similar goodness of the fit. The spectral properties of the tails in the two states are, however, quite different. This might indicate that the emission originates from different media in the two cases. Our results are compatible with a compact jet origin in the hard state and hybrid Comptonization in the soft state.

F. Cangemi↗

Potential origin of the state-dependent high-energy tail in the black hole microquasar Cygnus X-1 as seen with INTEGRAL

Context. 0.1–10 MeV observations of the black hole microquasar Cygnus X-1 have shown the presence of a spectral feature in the form of a power law in addition to the standard black body (0.1–10 keV) and Comptonization (10–200 keV) components observed by INTEGRAL in several black-hole X-ray binaries. This so-called “high-energy tail” was recently shown to be strong in the hard spectral state of Cygnus X-1, and, in this system, has been interpreted as the high-energy part of the emission from a compact jet. Aims. This result was nevertheless obtained from a data set largely dominated by hard state observations. In the soft state, only upper limits on the presence and hence the potential parameters of a high-energy tail could be derived. Using an extended data set, we aim to obtain better constraints on the properties of this spectral component in both states. Methods. We make use of data obtained from about 15 years of observations with the INTEGRAL satellite. The data set is separated into the different states and we analyze stacked state-resolved spectra obtained from the X-ray monitors, the gamma-ray imager, and the gamma-ray spectrometer (SPI) onboard. Results. A high-energy component is detected in both states, confirming its earlier detection in the hard state and its suspected presence in the soft state with INTEGRAL, as seen in a much smaller SPI data set. We first characterize the high-energy tail components in the two states through a model-independent, phenomenological analysis. We then apply physical models based on hybrid Comptonization (eqpair and belm). The spectra are well modeled in all cases, with a similar goodness of the fits. While in the semi-phenomenological approach the high-energy tail has similar indices in both states, the fits with the physical models seem to indicate slightly different properties. Based on this approach, we discuss the potential origins of the high-energy components in both the soft and hard states, and favor an interpretation where the high-energy component is due to a compact jet in the hard state and hybrid Comptonization in either a magnetized or nonmagnetized corona in the soft state.

accretion↗

Adaptive Machine Learning for Robust Diagnostics and Control of Time-Varying Particle Accelerator Components and Beams

Machine learning (ML) is growing in popularity for various particle accelerator applications including anomaly detection such as faulty beam position monitor or RF fault identification, for non-invasive diagnostics, and for creating surrogate models. ML methods such as neural networks (NN) are useful because they can learn input-output relationships in large complex systems based on large data sets. Once they are trained, methods such as NNs give instant predictions of complex phenomenon, which makes their use as surrogate models especially appealing for speeding up large parameter space searches which otherwise require computationally expensive simulations. However, quickly time varying systems are challenging for ML-based approaches because the actual system dynamics quickly drifts away from the description provided by any fixed data set, degrading the predictive power of any ML method, and limits their applicability for real time feedback control of quickly time-varying accelerator components and beams. In contrast to ML methods, adaptive model-independent feedback algorithms are by design robust to un-modeled changes and disturbances in dynamic systems, but are usually local in nature and susceptible to local extrema. In this work, we propose that the combination of adaptive feedback and machine learning, adaptive machine learning (AML), is a way to combine the global feature learning power of ML methods such as deep neural networks with the robustness of model-independent control. We present an overview of several ML and adaptive control methods, their strengths and limitations, and an overview of AML approaches.

97 MATHEMATICS AND COMPUTING↗

Analysis of Alternative Architectures for Cargo Lunar Landers

NASA’s Human Landing System (HLS) program has been working with commercial partners to develop human-class lunar landers to return the first American woman and next American man to the lunar surface in the mid 2020’s. In an effort to expand human presence beyond low Earth orbit, NASA’s Artemis program aims to facilitate a sustainable, long-term human presence in cis-lunar space. A component of this will require significant infrastructure to be delivered to the lunar surface. Delivering this infrastructure will require a significant lander capability that has yet to be developed. A thorough understanding of cargo lunar lander architectures is required such that select alternatives can be identified that best support the Artemis program’s objective of sustainability. The goal of this study is to aid NASA and its partners in the understanding of the cargo lunar lander trades space, as well as identify potential robust alternatives. The results will support NASA as it moves forward with key activities such as requirements formulation, agency strategic planning, and potential cargo lunar lander procurements. The study builds off of recent work performed by the Human Landing System program’s Architecture and Systems Analysis group to encompass a broad trade space of cargo lunar lander architecture alternatives. The current trade space as depicted by the morphological matrix and mission graph in Fig. 1 and Fig. 2, respectively, includes key alternative options that have become highly relevant due to current HLS activities and include on-orbit refueling, active cryogenic fluid management, Earth orbit aggregation, and global lunar access. The authors believe that there is also a statistically relevant impact of lander-payload configuration on the primary structure of the vehicle that could greatly impact alternative selection. Because of this, several conceptual lander-payload configurations will be evaluated to determine the level of impact. The current set of conceptual configurations are shown in Fig. 3 and Fig. 4. To aid the conceptual evaluation of these configurations, a catalogue of notional payloads has been developed that represent a wide range of masses and volumes that are expected to be delivered in support of a sustained human lunar presence, including pressurized and unpressurized rovers, surface habitats, power systems, and other support infrastructure. In order to execute this study in a timely fashion, a similar approach to that utilized in a similar 2019 study focused on 2024 human lunar sorties will be employed [1]. The team utilized a novel architecture synthesis framework currently being developed by NASA/MSFC to evaluate over 600,000 lunar lander architectures over a two month time frame [2]. From this large data set, varying ground rules and assumptions were applied as filters to explore the trade space to identify alternatives which exhibited robustness, as measured by launch vehicle payload margin, to absorb the natural growth that occurs during design maturation. The set of Earth-Moon system Delta-Vs assumed from the 2019 study, shown in Fig. 5, will be repurposed to accelerate model formulation for this effort. Additionally, current efforts in collaboration with the Georgia Institute of Technology’s Aerospace System Design Lab will be integrated to provide probabilistic modeling of the cargo lunar lander architectures to aid in identifying robust design alternatives [3]. The approach will help minimize potential impacts due to large levels of uncertainty inherent to pre phase-A conceptual design. By leveraging these past and present studies and partnerships, a highly detailed set of data can be generated in a short time period to aid NASA in the coming years to support the goal of a sustained human lunar presence.

Architectures↗

Training Ultrasound Image Classification Deep-Learning Algorithms for Pneumothorax Detection Using a Synthetic Tissue Phantom Apparatus

Ultrasound (US) imaging is a critical tool in emergency and military medicine because of its portability and immediate nature. However, proper image interpretation requires skill, limiting its utility in remote applications for conditions such as pneumothorax (PTX) which requires rapid intervention. Artificial intelligence has the potential to automate ultrasound image analysis for various pathophysiological conditions. Training models require large data sets and a means of troubleshooting in real-time for ultrasound integration deployment, and they also require large animal models or clinical testing. Here, we detail the development of a dynamic synthetic tissue phantom model for PTX and its use in training image classification algorithms. The model comprises a synthetic gelatin phantom cast in a custom 3D-printed rib mold and a lung mimicking phantom. When compared to PTX images acquired in swine, images from the phantom were similar in both PTX negative and positive mimicking scenarios. We then used a deep learning image classification algorithm, which we previously developed for shrapnel detection, to accurately predict the presence of PTX in swine images by only training on phantom image sets, highlighting the utility for a tissue phantom for AI applications.

Boice, Emily N. (ORCID:0000000171802842)↗

Interactive access and management for four-dimensional environmental data sets using McIDAS

This grant has fundamentally changed the way that meteorologists look at the output of their atmospheric models, through the development and wide distribution of the Vis5D system. The Vis5D system is also gaining acceptance among oceanographers and atmospheric chemists. Vis5D gives these scientists an interactive three-dimensional movie of their very large data sets that they can use to understand physical mechanisms and to trace problems to their sources. This grant has also helped to define the future direction of scientific visualization through the development of the VisAD system and its lattice data model. The VisAD system can be used to interactively steer and visualize scientific computations. A key element of this capability is the flexibility of the system's data model to adapt to a wide variety of scientific data, including the integration of several forms of scientific metadata.

Hibbard, William L.↗