Engineering PapersSearch

SEARCH · Engineering Papers

Results for “data augmentation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

A 360-degree and -order model of Venus topography

This report presents the most recent spherical harmonic topography model of Venus developed at Jet Propulsion Laboratory. It was produced by a spherical harmonic analysis of the most complete set of Magellan altimetry data, augmented by Pioneer Venus and Venera data. The harmonic coefficients of the topography were computed to degree and order 360. Compared to previous topography models, this one has the highest correlation with the gravity field of Venus.

Rappaport, Nicole

A 360-Degree and -Order Model of Venus Topography

This report presents the most recent spherical harmonic topography model of Venus developed at Jet Propulsion Laboratory. It was produced by a spherical harmonic analysis of the most complete set of Magellan altimetry data, augmented by Pioneer Venus and Venera data. The harmonic coefficients of the topography were computed to degree and order 360. Compared to previous topography models, this one has the highest correlation with the gravity field of Venus.

Venus venus topgraphy harmonic analysis spherical

NASA Pilot-Engaged Expert Response Using IBM Watson Technology: Prototype Evaluation of Knowledge Retrieval System

NASA Langley Research Center and IBM have been investigating the use of IBM Watson technology in aerospace research and development. One application of Watson technology is the Pilot-Engaged Expert Response (PEER) use case. The PEER system is envisioned as an in-cockpit advisor that will act as a source of situationally-relevant information for pilots and other flight crew members to assist in decision making about real-time events and situations that arise in the course of aircraft operations. PEER will make available vast stores of knowledge and information quickly and directly, putting important informational resources where they are needed most. IBM has worked with NASA to develop an architecture and articulate a roadmap for the development of the PEER system. That vision is built around Watson Discovery Advisor (WDA) software solution, derived from IBM's Jeopardy!-winning automatic question answering system. PEER makes use of WDA's sophisticated question-answering capabilities as its core, adding important User Interface components and other customizations for the cockpit environment, including communication with flight systems and other external data sources. The development plan for PEER includes four development stages, with the current project constituting the first phase. In this project, a prototype instance of PEER was successfully adapted to the aviation domain, enabling users to ask questions about aviation topics and receive useful and accurate answers to these questions. Major tasks accomplished include the development of procedures for domain adaptation through automatic lexicon extraction from domain glossaries; generation of question-answer training data which was used to train the system; and assessment of the effectiveness of domain adaptation, which showed a dramatic improvement in the ability of the PEER system to answer domain-relevant questions. In addition, the vision for the PEER system was pushed forward by the articulation of a plan for the automatic enhancement of question-answering with contextual information. This initial phase focused on two main goals: 1) the targeted domain adaptation of the underlying WDA system to the aviation domain; and, 2) the design of the software systems needed to leverage flight-contextual data. Domain adaptation of the WDA system proceeds via three main activities: Domain data ingestion, lexical customization and model training. A textual corpus consisting of 1,147 individual documents with more than 7.5 million words of text was ingested into the system and this served as the basis of all further development. A domain lexicon of over 3,500 aviation-domain terms was semi-automatically generated from domain documents and used to train the system. In addition, a set of over 500 question-answer (QA) pairs relevant to the PEER use case was developed; these were used to train and assess the system. These important first steps established the basis for the PEER system. In addition, steps were taken towards the integration of the PEER system into the cockpit environment with the development of a functional design for the Contextual Data Augmentation (CDA) subsystem. This subsystem brings to bear contextual data to improve system responses. It has three main submodules: the Contextual Data Collection module, the Contextual Data Selection module, and the Contextual QA Augmentation module. These modules form a processing pipeline that addresses the problems associated with automatically integrating information from external resources into the knowledge-retrieval mechanism.

Machine learning

A Simple Model of Pulsed Ejector Thrust Augmentation

A simple model of thrust augmentation from a pulsed source is described. In the model it is assumed that the flow into the ejector is quasi-steady, and can be calculated using potential flow techniques. The velocity of the flow is related to the speed of the starting vortex ring formed by the jet. The vortex ring properties are obtained from the slug model, knowing the jet diameter, speed and slug length. The model, when combined with experimental results, predicts an optimum ejector radius for thrust augmentation. Data on pulsed ejector performance for comparison with the model was obtained using a shrouded Hartmann-Sprenger tube as the pulsed jet source. A statistical experiment, in which ejector length, diameter, and nose radius were independent parameters, was performed at four different frequencies. These frequencies corresponded to four different slug length to diameter ratios, two below cut-off, and two above. Comparison of the model with the experimental data showed reasonable agreement. Maximum pulsed thrust augmentation is shown to occur for a pulsed source with slug length to diameter ratio equal to the cut-off value.

Wilson, Jack

Information Fusion and Data Analytics for Human Lunar Exploration (CIF REPORT: Detailed PI Write-up)

The Information Fusion & Data Analytics (IFDA) project commenced in FY20, continued through FY21, and its final platform development phase continues in FY22. The objective remains the fusion and rapid accessibility of large quantities of disparate sourced human spaceflight data. IFDA is a platform tailored for NA (S&MA) to develop highly advanced operational data integration and analysis techniques. IFDA leverages the JSC ER7 modeling, simulation,and data fusion capabilities to collect, warehouse, and augment data human exploration data integration and analysis techniques. The IFDA project’s integrated data visualizations have been demonstrated in two validation scenarios in FY21, and provided the architecture and platform basis for development of a full-scale data analysis suite and storage solution useful to all JSC organizations engaged in real time operations and safety tasks. Scenarioand prototypical development including the construction of a full scale data analysis suite and storage solution, useful to all JSC organizations engaged in real time operations and safety tasks, is central to IFDA Phase 3 and provides a demonstrable pathway for the Digital Transformation Program. IFDA Phase 3 is focused on data provider, data utilizer, and SME hands-on workshops that will conclude the Dem / Valphase and deliver a program-ready data integration tool as a product.

information fusion

Salvaging Data Records with Missing Data: Data Imputation using the Multivariate t Distribution

When doing multivariate data analysis, one commonobstacle is the presence of incomplete observations, i.e., observationsfor which one or more key fields are blank. Missing datais often countered by deleting entire observations that containmissing data. The negative effects of deleting entire observationsare multiple: deleting observations reduces sample size andcan also result in biased inferences even if data is missing atrandom. In addition, knowledge contained within incompleteobservations is knowledge lost when they are deleted– and theeffort spent collecting that knowledge is effort wasted. Data imputationmethods, or methods of statistically “filling-in” missingdata, can help combat small sample sizes by using the existinginformation in partially complete observations with the end goalof producing less biased and higher confidence inferences. Whena sample from a multivariate normal population is only partiallycomplete, and the missing data meets appropriate assumptions(missing at random), robust data imputation of the missing datacan be implemented with monotone data augmentation (MDA)using the multivariate t distribution.Missing data imputation is applied to data from the NASA InstrumentCost Model (NICM) using the MDA algorithm underthe assumption of having a multivariate t distribution with fixeddegrees of freedom. A sensitivity analysis to the degrees offreedom parameter is presented to demonstrate robustness ofthe multivariate t distribution when dealing with small samplesas compared to the multivariate normal distribution.

DiNicola, Michael

Information Fusion & Analytics for Human Lunar Exploration

The Information Fusion & Data Analytics (IFDA) project commenced in FY20, continued through FY21, and its final platform development phase continues in FY22. The objective remains the fusion and rapid accessibility of large quantities of disparate sourced human spaceflight data. IFDA is a platform tailored for NA (S&MA) to develop highly advanced operational data integration and analysis techniques. IFDA leverages the JSC ER7 modeling, simulation,and data fusion capabilities to collect, warehouse, and augment data human exploration data integration and analysis techniques. The IFDA project’s integrated data visualizations have been demonstrated in two validation scenarios in FY21, and provided the architecture and platform basis for development of a full-scale data analysis suite and storage solution useful to all JSC organizations engaged in real time operations and safety tasks. Scenarioand prototypical development including the construction of a full scale data analysis suite and storage solution, useful to all JSC organizations engaged in real time operations and safety tasks, is central to IFDA Phase 3 and provides a demonstrable pathway for the Digital Transformation Program. IFDA Phase 3 is focused on data provider, data utilizer, and SME hands-on workshops that will conclude the Dem / Valphase and deliver a program-ready data integration tool as a product.

information fusion

Extracting Material Property Measurements from Scientific Literature with Limited Annotations

Extracting material property data from scientific text is pivotal for advancing data-driven research in chemistry and materials science; however, the extensive annotation effort required to produce training data for named entity recognition (NER) models for this task often makes it a barrier to extracting specialized data sets. Here, in this work, we present a comparative study of the conventional, supervised NER methodology to alternative few-shot learning architectures and large language model (LLM)-based approaches that mitigate the need to label large training data sets. We find that the best-performing LLM (GPT-4o) not only excels in directly extracting relevant material properties based on limited examples but also enhances supervised learning through data augmentation. We supplement our findings with error and data quality assessments to provide a nuanced understanding of factors that impact property measurement extraction.

36 MATERIALS SCIENCE

The report of the Gravity Field Workshop

A Gravity Field Workshop was convened to review the actions which could be taken prior to a GRAVSAT mission to improve the Earth's gravity field model. This review focused on the potential improvements in the Earth's gravity field which could be obtained using the current satellite and surface gravity data base. In particular, actions to improve the quality of the gravity field determination through refined measurement corrections, selected data augmentation and a more accurate reprocessing of the data were considered. In addition, recommendations were formulated which define actions which NASA should take to develop the necessary theoretical and computation techniques for gravity model determination and to use these approaches to improve the accuracy of the Earth's gravity model.

Smith, D. E.

Comparison of model and flight test data for an augmented jet flap STOL research aircraft

Aerodynamic design data for the Augmented Jet Flap STOL Research Aircraft or commonly known as the Augmentor-Wing Jet-STOL Research Aircraft was based on results of tests carried out on a large scale research model in the NASA Ames 40- by 80-Foot Wind Tunnel. Since the model differs in some respects from the aircraft, precise correlation between tunnel and flight test is not expected, however the major areas of confidence derived from the wind tunnel tests are delineated, and for the most part, tunnel results compare favorably with flight experience. In some areas the model tests were known to be nonrepresentative so that a degree of uncertainty remained: these areas of greater uncertainty are identified, and discussed in the light of subsequent flight tests.

Cook, W. L.

HST WFC3 Early Release Science: Emission-Line Galaxies from IR Grism Observations

We present grism spectra of emission line galaxies (ELGs) from 0.6-1.6 microns from the Wide Field Camera 3 (WFC3) on the Hubble Space Telescope (HST). These new infrared grism data augment previous optical Advanced Camera for Surveys G800L (0.6-0.95 micron) grism data in GOODS South, extending the wavelength coverage well past the G800L red cutoff. The ERS grism field was observed at a depth of 2 orbits per grism, yielding spectra of hundreds of faint objects, a subset of which are presented here. ELGs are studied via the Ha, [O III ], and [OII] emission lines detected in the redshift ranges 0.2 less than or equal to z less than or equal to 1.6, 1.2 less than or equal to z less than or equal to 2.4 and 2.0 less than or equal to z less than or equal to 3.6 respectively in the G102 (0.8-1.1 microns; R approximately 210) and C141 (1.1-1.6 microns; R approximately 130) grisms. The higher spectral resolution afforded by the WFC3 grisms also reveals emission lines not detectable with the G800L grism (e.g., [S II] and [S III] lines). From these relatively shallow observations, line luminosities, star formation rates, and grism spectroscopic redshifts are determined for a total of 25 ELGs to M(sub AB)(F098M) approximately 25 mag. The faintest source in our sample with a strong but unidentified emission line--is MAB(F098M)=26.9 mag. We also detect the expected trend of lower specific star formation rates for the highest mass galaxies in the sample, indicative of downsizing and discovered previously from large surveys. These results demonstrate the remarkable efficiency and capability of the WFC3 NIR grisms for measuring galaxy properties to faint magnitudes.

Straughn, A. N.

CIF Report - Information Fusion and Data Analytics for Human Lunar Exploration

This project leverages the Concept Exploration Laboratory (CEL) to collect, warehouse, and augment data relevant to human lunar exploration as a platform for NA (S&MA) to develop operational data integration techniques. The project capitalizes on 16+ years of CEL experience applied to NASA, DoD, the City of Houston, the State of Texas, and private industry. The integrated data will be utilized in the two scenarios described in a definition of concept for development of a full scale data analysis suite and storage solution, useful to all JSC organizations engaged in real time operations and safety tasks, and may be useful as pathfinders for the Digital Transformation Program.

information fusion

Describing Point Defect Topology in 2D Energy Materials Through Computer Vision

Point defects such as vacancies and impurity atoms strongly impact the performance of 2D materials. Traditional efforts often rely on manual detection, a process that is time-intensive, prone to human error, and challenging to scale. Here we leverage machine learning (ML) methods to identify and quantify vacancies within 2D transition metal carbides (Ti3C2, MXenes), aiming to expedite detection while improving accuracy. MXenes exhibit valuable defect-defined electrochemical properties, but we currently lack statistical understanding of defect topology needed to fully harness these materials. Here we employ a convolutional neural network for semantic segmentation of experimental MXene images, opening an opportunity to conduct a rigorous statistical study on defect hierarchy while investigating local relaxation in the lattice. We show how the integration of ML can yield fundamental insight into point defects, providing a powerful tool that will play an increasingly crucial role in the future of materials science. ML is often not just a matter of straightforward application, and pretrained models proved ineffective in this case. Instead, we trained our own neural network (NN) and applied data augmentation techniques and fine-tuning to the training dataset. Since labeled microscopy data is often scarce, we developed training data from a previously published wide-frame MXene image, using customized Gaussian fitting to locate atomic positions. Our trained model was then applied to a large dataset of experimental images, enabling a statistical study of defect configurations across three samples prepared with different HF etchant concentrations (5%, 9.1%, and 12.5%), as shown in Fig. 1. This also allowed us to investigate local strain around vacancies, though we find that we are limited by the precision of measurements using high-angle annular dark field (HAADF) images, as shown in Fig. 2. This study demonstrates how ML enables large-scale, quantitative analysis of atomic defects - an otherwise infeasible task with traditional methods. While our NN was specialized for Ti3C2 MXenes, the pipeline we developed provides a foundation for future ML models tailored to other materials. Ultimately, we envision embedding the NN onto the microscope to give real-time feedback to the user. To make this a reality, continued work is necessary to fully understand the NN's capabilities and limitations. This study gets one step closer to our goals of automated experimentation moving away from traditional methods of manual labeling. As ML capabilities advance, we hope to continue adapting and applying these techniques in microscopy.

2D materials

2025 TEM Workshop

The TEM Data Management Workshop will take place on August 26 from 9 a.m. to 12 p.m. MT, and will be held virtually on TEAMS. The primary goal of this workshop is to engage NSUF users and stakeholders in discussions about the data needs for the utilization of AI and ML in the analysis of TEM data. Key topics to be covered include data storage, data sharing, data tagging, metadata inclusion, standardized data formats, data augmentation, and annotated training datasets. Additionally, the workshop will provide valuable insights into resources such as the Nuclear Research Data System (NRDS) for data storage and sharing, as well as open-source codes for data analysis.

Bachhav, Mukesh

Detecting change as it occurs

Traditionally climate changes have been detected from long series of observations and long after they have happened. Our 'inverse sequential' procedure, for detecting change as soon as it occurs, describes the existing or most recent data by their frequency distribution. Its parameter(s) are estimated both from the existing set of observations and from the same set augmented by 1,2,....j new observations. Individual-value probability products ('likelihoods') are used to form ratios which yield two probabilities for erroneously accepting the existing parameter(s) as valid for the augmented data set, and vice versa. A genuine parameter change is signaled when these probabilities (or a more stable compound probability) show a progressive decrease. New parameter values can then be estimated from the new observations alone using standard statistical techniques. The inverse sequential procedure will be illustrated for global annual mean temperatures (assumed normally distributed), and for annual numbers of North Atlantic hurricanes (assumed to represent Poisson distributions). The procedure was developed, but not yet tested, for linear or exponential trends, and for chi-squared means or degrees of freedom, a special measure of autocorrelation.

Radok, Uwe

Transplatformer: translating toxicogenomic profiles between generations of platforms

Background Transcriptomic profiling technologies have advanced the analysis of biological and toxicological responses. However, substantial differences in probe design, dynamic range, gene coverage, and preprocessing pipelines across platforms introduce artifacts that limit cross-study integration and hinder the reuse of historical datasets. We aim to develop computational methods for accurate cross-platform translation to maximize the value of legacy resources. Results We present TransPlatformer a deep learning framework for translating gene expression profiles across heterogeneous toxicogenomics platforms. TransPlatformer employs a novel attention-based architecture to map high-dimensional fold-change vectors from legacy microarray technologies to current platforms. Models are trained and evaluated using DrugMatrix, spanning three technological generations. We investigate mixed-tissue, single-tissue, and cross-tissue training paradigms and benchmark performance against multilayer perceptron and matrix-completion baselines. In mixed-tissue training, TransPlatformer achieves a greater than 50% reduction in mean absolute error (0.043 vs. 0.09) and nearly doubles Pearson correlation ( ≈ 0.71 vs. 0.37) relative to baseline methods. Importantly, TransPlatformer preserves rare but biologically meaningful over- and under-expressed signals, with mean absolute error below 0.22. Single-tissue models yield further improvements for well-represented organs, such as a 10% reduction in liver mean absolute error, while underscoring the need for data augmentation strategies in low-sample tissues.ra Conclusions TransPlatformer provides an effective and scalable computational solution for cross-platform transcriptomic translation. By enabling biologically faithful harmonization of gene expression data, the proposed approach facilitates the reuse of legacy toxicogenomics datasets, enhances downstream biomarker discovery, and supports more reproducible predictive modeling in toxicology.

59 BASIC BIOLOGICAL SCIENCES

Roadmap for transforming heterogeneous catalysis with artificial intelligence

Artificial intelligence (AI) is poised to transform heterogeneous catalysis, opening avenues for catalytic materials discovery. By uncovering intricate patterns in high-dimensional data, AI has been reshaping our pursuit of sustainable catalytic processes across the energy, environmental and chemical sectors. This promise, however, hinges on overcoming fundamental barriers, including limitations in data availability and quality, challenges in the generalizability and interpretability of data-augmented decisions, and the persistent gap between in silico predictions and experiments. Furthermore, we outline a forward-looking roadmap for deeply integrating AI into heterogeneous catalysis with an AI-ready data ecosystem, multimodal foundation models, and ultimately autonomous laboratories to accelerate the development of next-generation catalytic technologies via AI-empowered human–machine collaboration.

Computational methods

Names Don't Fly: Smart Filters for Profanity Detection and Classification in User-Generated Content

Generally, names associate with a person’s identity. But what if in the pretext of a legitimate name and given the opportunity, users of software provide names to online web forms that carry along offensive language, slurs, and other profanity that is then sent to Mars ? The answer is simple: they don’t fly. In this paper,we perform model explorations to detect and classify inappropriate content in the names submitted from people across the world to ‘Send Your Names to MARS’ public engagement campaign.We propose a novel pipeline approach, that can effectively overcome the issues of lack of negative samples, noisy labels by gathering expert knowledge over time with human(s) in the loop and data augmentation, and achieve high accuracy in classifying inappropriate names with very little or no context. We describe cloud-based infrastructure to deploy our application and run predictions on large-scale data through our pipeline and achieve significant speedup over offline processes, with enhanced reliability and security.

Soderstrom, Tomas