Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data processing automation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15

Eukaryotic genomes from a global metagenomic data set illuminate trophic modes and biogeography of ocean plankton

ABSTRACT Metagenomics is a powerful method for interpreting the ecological roles and physiological capabilities of mixed microbial communities. Yet, many tools for processing metagenomic data are neither designed to consider eukaryotes nor are they built for an increasing amount of sequence data. EukHeist is an automated pipeline to retrieve eukaryotic and prokaryotic metagenome-assembled genomes (MAGs) from large-scale metagenomic sequence data sets. We developed the EukHeist workflow to specifically process large amounts of both metagenomic and/or metatranscriptomic sequence data in an automated and reproducible fashion. Here, we applied EukHeist to the large-size fraction data (0.8–2,000 µm) from Tara Oceans to recover both eukaryotic and prokaryotic MAGs, which we refer to as TOPAZ (Tara Oceans Particle-Associated MAGs). The TOPAZ MAGs consisted of >900 environmentally relevant eukaryotic MAGs and >4,000 bacterial and archaeal MAGs. The bacterial and archaeal TOPAZ MAGs expand upon the phylogenetic diversity of likely particle- and host-associated taxa. We use these MAGs to demonstrate an approach to infer the putative trophic mode of the recovered eukaryotic MAGs. We also identify ecological cohorts of co-occurring MAGs, which are driven by specific environmental factors and putative host-microbe associations. These data together add to a number of growing resources of environmentally relevant eukaryotic genomic information. Complementary and expanded databases of MAGs, such as those provided through scalable pipelines like EukHeist, stand to advance our understanding of eukaryotic diversity through increased coverage of genomic representatives across the tree of life. IMPORTANCE Single-celled eukaryotes play ecologically significant roles in the marine environment, yet fundamental questions about their biodiversity, ecological function, and interactions remain. Environmental sequencing enables researchers to document naturally occurring protistan communities, without culturing bias, yet metagenomic and metatranscriptomic sequencing approaches cannot separate individual species from communities. To more completely capture the genomic content of mixed protistan populations, we can create bins of sequences that represent the same organism (metagenome-assembled genomes [MAGs]). We developed the EukHeist pipeline, which automates the binning of population-level eukaryotic and prokaryotic genomes from metagenomic reads. We show exciting insight into what protistan communities are present and their trophic roles in the ocean. Scalable computational tools, like EukHeist, may accelerate the identification of meaningful genetic signatures from large data sets and complement researchers’ efforts to leverage MAG databases for addressing ecological questions, resolving evolutionary relationships, and discovering potentially novel biodiversity.

59 BASIC BIOLOGICAL SCIENCES↗

Protein Crystal Growth

In order to rapidly and efficiently grow crystals, tools were needed to automatically identify and analyze the growing process of protein crystals. To meet this need, Diversified Scientific, Inc. (DSI), with the support of a Small Business Innovation Research (SBIR) contract from NASA s Marshall Space Flight Center, developed CrystalScore(trademark), the first automated image acquisition, analysis, and archiving system designed specifically for the macromolecular crystal growing community. It offers automated hardware control, image and data archiving, image processing, a searchable database, and surface plotting of experimental data. CrystalScore is currently being used by numerous pharmaceutical companies and academic and nonprofit research centers. DSI, located in Birmingham, Alabama, was awarded the patent Method for acquiring, storing, and analyzing crystal images on March 4, 2003. Another DSI product made possible by Marshall SBIR funding is VaporPro(trademark), a unique, comprehensive system that allows for the automated control of vapor diffusion for crystallization experiments.

Source record↗

Investigation into Cloud Computing for More Robust Automated Bulk Image Geoprocessing

Geospatial resource assessments frequently require timely geospatial data processing that involves large multivariate remote sensing data sets. In particular, for disasters, response requires rapid access to large data volumes, substantial storage space and high performance processing capability. The processing and distribution of this data into usable information products requires a processing pipeline that can efficiently manage the required storage, computing utilities, and data handling requirements. In recent years, with the availability of cloud computing technology, cloud processing platforms have made available a powerful new computing infrastructure resource that can meet this need. To assess the utility of this resource, this project investigates cloud computing platforms for bulk, automated geoprocessing capabilities with respect to data handling and application development requirements. This presentation is of work being conducted by Applied Sciences Program Office at NASA-Stennis Space Center. A prototypical set of image manipulation and transformation processes that incorporate sample Unmanned Airborne System data were developed to create value-added products and tested for implementation on the "cloud". This project outlines the steps involved in creating and testing of open source software developed process code on a local prototype platform, and then transitioning this code with associated environment requirements into an analogous, but memory and processor enhanced cloud platform. A data processing cloud was used to store both standard digital camera panchromatic and multi-band image data, which were subsequently subjected to standard image processing functions such as NDVI (Normalized Difference Vegetation Index), NDMI (Normalized Difference Moisture Index), band stacking, reprojection, and other similar type data processes. Cloud infrastructure service providers were evaluated by taking these locally tested processing functions, and then applying them to a given cloud-enabled infrastructure to assesses and compare environment setup options and enabled technologies. This project reviews findings that were observed when cloud platforms were evaluated for bulk geoprocessing capabilities based on data handling and application development requirements.

Brown, Richard B.↗

Fostering Geothermal Machine Learning Success: Elevating Big Data Accessibility and Automated Data Standardization in the Geothermal Data Repository

The Department of Energy's (DOE) Geothermal Data Repository (GDR) has implemented improvements to both its data lakes and its data standards and automated data pipelines. The GDR data lakes have reduced storage and compute-related barriers to using large geothermal datasets, enabling these large datasets to be accessed by anyone with a modern computer and internet access. More recently, the GDR has been working to further reduce barriers through streamlining the data intake process, educating users on the process and requirements, and aiding users in accessing data from the data lakes. These improvements have augmented the quantity of datasets the GDR is able to accept into its data lakes and have enabled users who are new to cloud tools to access these datasets more easily, overall increasing the accessibility of big geothermal data for use in machine learning and other projects. In addition, the GDR now has built-in data standards and pipelines for drilling data, geospatial data, and distributed acoustic sensing (DAS) data. These standardization efforts aim to enhance the real-world applicability of geothermal machine learning outcomes by improving the quality of training data. Specifically, through standardizing high-value datasets, the GDR is reducing project-specific data curation requirements, thus allowing more time for actual research. By automating this process, the burden of standardization is lifted from the user, ultimately increasing the availability of standardized data.

15 GEOTHERMAL ENERGY↗

Fostering Geothermal Machine Learning Success: Elevating Big Data Accessibility and Automated Data Standardization in the Geothermal Data Repository: Preprint

The Department of Energy's (DOE) Geothermal Data Repository (GDR) has implemented improvements to both its data lakes and its data standards and automated data pipelines. The GDR data lakes have reduced storage and compute-related barriers to using large geothermal datasets, enabling these large datasets to be accessed by anyone with a modern computer and internet access. More recently, the GDR has been working to further reduce barriers through streamlining the data intake process, educating users on the process and requirements, and aiding users in accessing data from the data lakes. These improvements have augmented the quantity of datasets the GDR is able to accept into its data lakes and have enabled users who are new to cloud tools to access these datasets more easily, overall increasing the accessibility of big geothermal data for use in machine learning and other projects. In addition, the GDR now has built-in data standards and pipelines for drilling data, geospatial data, and distributed acoustic sensing (DAS) data. These standardization efforts aim to enhance the real-world applicability of geothermal machine learning outcomes by improving the quality of training data. Specifically, through standardizing high-value datasets, the GDR is reducing project-specific data curation requirements, thus allowing more time for actual research. By automating this process, the burden of standardization is lifted from the user, ultimately increasing the availability of standardized data.

accessibility↗

Artificial Intelligence in Nuclear Physics

Artificial Intelligence (AI) and Machine Learning (ML) are rapidly developing fields providing data-driven algorithms to predict, classify, and make decisions based on data. Nuclear Physics Research is data-driven and AI/ML techniques have been implemented for experiment and accelerator control, in theoretical applications, and in data processing and analysis. These algorithms open possibilities for automation, thereby augmenting human capabilities. Additionally, Open Science is enabled by simultaneous analyses of multiple data sources, leading to scientific knowledge. This talk will summarize current applications of AI/ML in nuclear physics, as well as accelerator applications, and will cover upcoming initiatives and research in AI/ML.

Jeske, Torri↗

Systems and methods for control of polymer reactions and processing using automatic continuous online monitoring

Manual and automatic methods and devices using a ACOMP system for active control of polymerization reaction processes. An ideal desired trajectory of one or more reaction and polymer characteristics can be established to produce a desired final polymer product with specified characteristics from a polymerization reaction process. A current reaction trajectory of a polymerization reaction process can be driven to an ideal or desired reaction trajectory. In a manual embodiment an operator can use ACOMP data to adjust process variables in order to drive the current reaction trajectory toward the ideal or desired reaction trajectory. In an automated mode a control program can use ACOMP data to make adjustments to process variables to drive the polymerization reaction process toward the desired trajectory as closely as possible either empirically or by solving the governing equations for the polymerization reaction process.

Reed, Wayne Frederick↗

nmRanalysis: An Open-Source Web Application for Semi-automated NMR Metabolite Profiling

Though data acquisition and initial signal pre-processing of nuclear magnetic resonance (NMR) spectra have achieved high degrees of automation, downstream processing - specifically the profiling of spectra - has bottlenecked the overall NMR analysis workflow. Several efforts have been made to mitigate this bottleneck, but these solutions often trade an increase in automation for limitations elsewhere. Here, in this technical note, we introduce nmRanalysis, a user-friendly web-application that integrates the strengths of existing profiling tools for a more automated profiling workflow. nmRa-nalysis additionally incorporates novel features, including a machine-learning-driven recommender system for me-tabolite identification, further increasing the utility of nmRanalysis over the individual tools that it incorporates.

Flores, Javier E. [Pacific Northwest National Labo↗

Cinema:Snap: Real-time tools for analysis of dynamic diamond anvil cell experiment data

We report we developed tools and a workflow for real-time analysis of data from dynamic diamond anvil cell experiments performed at user light sources. These tools allow users to determine the phases of matter observed during the compression of materials in order to make decisions during an experiment to improve the quality of experimental results and maximize the use of scarce experimental facility time. The tools fill a gap in dynamic compression data analysis tools that are real-time, are flexible to the needs of high-pressure scientists, connect to automated processing of results, can be easily incorporated into workflows with existing tools and data formats, and support remote experimental data analysis workflows. Specific analytics developed include novel automated two-peak analysis for overlapping peaks and multiple phases, coordinated views of pressure and temperature values, full-compression contour plots, and configurable views of integrated x-ray diffraction. We present an experimental use case to show how the tools produce real-time analytics that help the scientists revise parameters for the next compression

47 OTHER INSTRUMENTATION↗

Determination of design and operation parameters for upper atmospheric research instrumentation to yield optimum resolution with deconvolution, appendix 2

This thesis reviews the technique established to clear channels in the Power Spectral Estimate by applying linear combinations of well known window functions to the autocorrelation function. The need for windowing the auto correlation function is due to the fact that the true auto correlation is not generally used to obtain the Power Spectral Estimate. When applied, the windows serve to reduce the effect that modifies the auto correlation by truncating the data and possibly the autocorrelation has on the Power Spectral Estimate. It has been shown in previous work that a single channel has been cleared, allowing for the detection of a small peak in the presence of a large peak in the Power Spectral Estimate. The utility of this method is dependent on the robustness of it on different input situations. We extend the analysis in this paper, to include clearing up to three channels. We examine the relative positions of the spikes to each other and also the effect of taking different percentages of lags of the auto correlation in the Power Spectral Estimate. This method could have application wherever the Power Spectrum is used. An example of this is beam forming for source location, where a small target can be located next to a large target. Other possibilities extend into seismic data processing. As the method becomes more automated other applications may present themselves.

Ioup, George E.↗

Description of Selected Algorithms and Implementation Details of a Concept-Demonstration Aircraft VOrtex Spacing System (AVOSS)

A ground-based system has been developed to demonstrate the feasibility of automating the process of collecting relevant weather data, predicting wake vortex behavior from a data base of aircraft, prescribing safe wake vortex spacing criteria, estimating system benefit, and comparing predicted and observed wake vortex behavior. This report describes many of the system algorithms, features, limitations, and lessons learned, as well as suggested system improvements. The system has demonstrated concept feasibility and the potential for airport benefit. Significant opportunities exist however for improved system robustness and optimization. A condensed version of the development lab book is provided along with samples of key input and output file types. This report is intended to document the technical development process and system architecture, and to augment archived internal documents that provide detailed descriptions of software and file formats.

Hinton, David A.↗

UH-60A Airloads Flight Test Program: Data Counter 9017

Aeromechanics Branch interns at Ames Research Center have been directly contributing to the data quality analysis and reporting of the UH-60A Airloads Flight Test Program for many years. In chronological order (together with the semester and year): Caroline Edwards (Summer 2011); Joni DeGuzman and Carson Turner (Fall 2011); Eric Fritz (Spring 2012); Connor Beierle (Fall 2012); Christopher Olinger (Spring and Summer 2013); Needa Lin, Anatole Levkoff, Maxwell Loebig, Jose Orejel, Megan Prout, and Albert Sue (Summer 2014); Jared Archey (Fall 2014); Alexander Crone (Summer 2015); Jeffrey Diament, Austin Djang, and Jessica Swan (Summer 2016); Makenzie Allen (Summer 2017); Colin Lauzon (Fall 2017); Eric Gilkey (Spring 2018); and Nicholas Masso (Summer 2019). These interns have spent their internships reviewing flight logs, extracting the data out of TRENDS, formatting the data into spreadsheets, writing code to automate the process, and plotting results. Without their efforts, much of the work would be unfinished. The authors appreciate the achievements of the UH-60A Airloads Working Group during its 20-year lifetime, as well the contributions of Randy Peterson, Tom Norman, and William Warmbrodt to the data processing and assistance with the report preparation. Lastly, this report is dedicated to William Bousman for his efforts preceding, during, and subsequent to the UH-60A Airloads Flight Test Program.

Kufeld, Robert M.↗

LaRC SmartLab Apps For Instrument Control and Data Processing: Laboratory Environment Monitor

The LaRC SmartLab applications are a series of software tools to greatly enhance researcher efficiency by streamlining and automating workflows. Python scripts and applications are increasingly being used in scientific workflows, including for instrument control and data processing. Interactive Python scripting environments such as JupyterLab provide powerful tools for using Python. In some use cases, the development of standalone applications with dedicated graphical user interfaces can enhance the utility of the code and open it up to more users, including non-programmers. Here, we describe a Python based application for communicating with, and displaying data from, iTHX Temperature, Humidity, and Dew Point probes. We discuss the set up and use of the application as well as its implementation. We also highlight the use of Simulated probes to enable users and developers to familiarize with or debug the application, even when they do not have access to the physical hardware in the laboratory.

LaRC SmartLab↗

Digital Radar-Signal Processors Implemented in FPGAs

High-performance digital electronic circuits for onboard processing of return signals in an airborne precipitation- measuring radar system have been implemented in commercially available field-programmable gate arrays (FPGAs). Previously, it was standard practice to downlink the radar-return data to a ground station for postprocessing a costly practice that prevents the nearly-real-time use of the data for automated targeting. In principle, the onboard processing could be performed by a system of about 20 personal- computer-type microprocessors; relative to such a system, the present FPGA-based processor is much smaller and consumes much less power. Alternatively, the onboard processing could be performed by an application-specific integrated circuit (ASIC), but in comparison with an ASIC implementation, the present FPGA implementation offers the advantages of (1) greater flexibility for research applications like the present one and (2) lower cost in the small production volumes typical of research applications. The generation and processing of signals in the airborne precipitation measuring radar system in question involves the following especially notable steps: The system utilizes a total of four channels two carrier frequencies and two polarizations at each frequency. The system uses pulse compression: that is, the transmitted pulse is spread out in time and the received echo of the pulse is processed with a matched filter to despread it. The return signal is band-limited and digitally demodulated to a complex baseband signal that, for each pulse, comprises a large number of samples. Each complex pair of samples (denoted a range gate in radar terminology) is associated with a numerical index that corresponds to a specific time offset from the beginning of the radar pulse, so that each such pair represents the energy reflected from a specific range. This energy and the average echo power are computed. The phase of each range bin is compared to the previous echo by complex conjugate multiplication to obtain the mean Doppler shift (and hence the mean and variance of the velocity of precipitation) of the echo at that range.

Berkun, Andrew↗

The Human Research Program Grant Lifecycle & Data Integration Schedule: Infographic

The Human Research Program (HRP) Grant Lifecycle process can be described using information from several government authoritative sources with overlapping generalizations. It is the responsibility of the HRP Program Planning & Control (PP&C) Office and the Data Management Integration Office (DMIO)to interpret this information into a cohesive process that can be communicated to the human research organization. PP&C and DMIO have been exploring training opportunities to facilitate understanding through the form of infographics. This communication tool combines eye-catching visuals and text making complex information and data more digestible and shareable. The infographic poster tells a short story about a federally awarded grant and is intended to help both the new and experienced HRP workforce understand the timeline for a research procurement and where their specific tasks fall within the 4 phases of the grant lifecycle. The Pre-Award, Award, Post-Award, and Closeout phases are overlayed with business swimlanes so HRP stakeholders can see where they fit into the timeline rather than working in an isolated part. This includes the Chief Scientist Office (CSO), PP&C, the Data Management Integration Office (DMIO), the Principal Investigator (PI), the Element stakeholders, the LSDA Archivists, and the Research & Operations Integration (ROI) team along with the stakeholders in the Human Health and Performance Directorate. Additionally, the poster introduces (1) the HRP Data Integration Schedule identified in the HRP Data Management Plan HRP-48047, Table 8-2 and (2) a Smartsheet Solution to automate the grants tracking business process. Emphasis is given to the Data Integration Schedule which is the basis for the milestone tasks that are to be completed during the Post-Award phase in the grant lifecycle. The poster is also an opportunity to present the Smartsheet Solution, a software collaboration and work management tool used to assign tasks and track project progress. This PP&C FY24 effort will facilitate an end-to-end solution to track grants and provide insight to the health of HRP grant research procurements using metrics, reports, and dashboards.

J Peace↗

Verification of a New NOAA/NSIDC Passive Microwave Sea-Ice Concentration Climate Record

A new satellite-based passive microwave sea-ice concentration product developed for the National Oceanic and Atmospheric Administration (NOAA)Climate Data Record (CDR) programme is evaluated via comparison with other passive microwave-derived estimates. The new product leverages two well-established concentration algorithms, known as the NASA Team and Bootstrap, both developed at and produced by the National Aeronautics and Space Administration (NASA) Goddard Space Flight Center (GSFC). The sea ice estimates compare well with similar GSFC products while also fulfilling all NOAA CDR initial operation capability (IOC) requirements, including (1) self describing file format, (2) ISO 19115-2 compliant collection-level metadata,(3) Climate and Forecast (CF) compliant file-level metadata, (4) grid-cell level metadata (data quality fields), (5) fully automated and reproducible processing and (6) open online access to full documentation with version control, including source code and an algorithm theoretical basic document. The primary limitations of the GSFC products are lack of metadata and use of untracked manual corrections to the output fields. Smaller differences occur from minor variations in processing methods by the National Snow and Ice Data Center (for the CDR fields) and NASA (for the GSFC fields). The CDR concentrations do have some differences from the constituent GSFC concentrations, but trends and variability are not substantially different.

Passive Microwave↗

Using Information Automation and Human Technology Integration to Implement Integrated Operations for Nuclear

The purpose of the research effort described in this report is to develop and demonstrate an approach to the design and implementation of advanced, automated systems intended to increase operational efficiencies at nuclear power plants. In particular, we describe methods for considering human-technology integration (HTI) issues throughout the various phases of system design, test, and implementation and how these considerations help promote effective design. This research project will develop planning tools and comprehensive guidance on how HTI principles and methods, in combination with information automation technologies, can enable effective data integration and coordination for full nuclear plant modernization. Specifically, this research project will develop an approach to automate the mapping of data from plant systems and processes to application needs, thereby significantly reducing the amount of human workload currently required for the execution of these tasks. In addition to developing an automated solution as a replacement for these tasks, this research will also analyze digitalization’s effectiveness in reducing operational costs of compliance related activities. Compliance activities are estimated to account for as much as 50% of operations and maintenance (i.e., non-fuel and non-capital) costs of plant operation.

97 MATHEMATICS AND COMPUTING↗

Machine learning for domain transfer between simulated and experimental 2D X-ray diffraction patterns using generative adversarial networks

X-ray diffraction (XRD) is a well-established technique for analyzing materials at an atomic level. Dynamic compression experiments (DCE), in which materials are subject to extreme pressures, can provide fundamental understanding to pressure-induced phase transitions and compression of the crystal lattice. The analysis of XRD patterns from highly compressed samples is non-trivial given the sparsity of data, high experimental costs, and the fact that the data is often marred with X-ray background and other artifacts. While accurate computational frameworks exist, they solve the forward problem—from structures and orientations to XRD patterns. Solving the inverse problem for 2D experimental diffraction patterns is currently a complex manual process of matching and comparing experimentally observed patterns to computationally generated ones. Machine learning is a promising tool for automating the matching process but often requires data-intensive architectures. Here, in this study, we use a CycleGAN to translate the domain of limited experimental data to a domain in which there is readily available simulated data. This domain shift allows data-intensive machine learning models that have only been trained on simulated XRD patterns to be used in the analysis of experiments.

Brozak, Samantha Jean [Sandia National Laboratorie↗