Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Functional Data Analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Real-space visualization of short-range antiferromagnetic correlations in a magnetically enhanced thermoelectric

Short-range magnetic correlations can significantly increase the thermopower of magnetic semiconductors, representing a noteworthy development in the decades-long effort to develop high-performance thermoelectric materials. Here, we reveal the nature of the thermopower-enhancing magnetic correlations in the antiferromagnetic semiconductor MnTe. Using magnetic pair distribution function analysis of neutron scattering data, we obtain a detailed, real-space view of robust, nanometer-scale, antiferromagnetic correlations that persist into the paramagnetic phase above the Neel temperature $T_N$ = 307 K. In this work, the magnetic correlation length in the paramagnetic state is significantly longer along the crystallographic c axis than within the ab plane, pointing to anisotropic magnetic interactions. Ab initio calculations of the spin-spin correlations using density functional theory in the disordered local moment approach reproduce this result with quantitative accuracy. These findings constitute the first real-space picture of short-range spin correlations in a magnetically enhanced thermoelectric and inform future efforts to optimize thermoelectric performance by magnetic means.

36 MATERIALS SCIENCE↗

Safe Upper-Bounds Inference of Energy Consumption for Java Bytecode Applications

Many space applications such as sensor networks, on-board satellite-based platforms, on-board vehicle monitoring systems, etc. handle large amounts of data and analysis of such data is often critical for the scientific mission. Transmitting such large amounts of data to the remote control station for analysis is usually too expensive for time-critical applications. Instead, modern space applications are increasingly relying on autonomous on-board data analysis. All these applications face many resource constraints. A key requirement is to minimize energy consumption. Several approaches have been developed for estimating the energy consumption of such applications (e.g. [3, 1]) based on measuring actual consumption at run-time for large sets of random inputs. However, this approach has the limitation that it is in general not possible to cover all possible inputs. Using formal techniques offers the potential for inferring safe energy consumption bounds, thus being specially interesting for space exploration and safety-critical systems. We have proposed and implemented a general frame- work for resource usage analysis of Java bytecode [2]. The user defines a set of resource(s) of interest to be tracked and some annotations that describe the cost of some elementary elements of the program for those resources. These values can be constants or, more generally, functions of the input data sizes. The analysis then statically derives an upper bound on the amount of those resources that the program as a whole will consume or provide, also as functions of the input data sizes. This article develops a novel application of the analysis of [2] to inferring safe upper bounds on the energy consumption of Java bytecode applications. We first use a resource model that describes the cost of each bytecode instruction in terms of the joules it consumes. With this resource model, we then generate energy consumption cost relations, which are then used to infer safe upper bounds. How energy consumption for each bytecode instruction is measured is beyond the scope of this paper. Instead, this paper is about how to infer safe energy consumption estimations assuming that those energy consumption costs are provided. For concreteness, we use a simplified version of an existing resource model [1] in which an energy consumption cost for individual Java opcodes is defined.

Navas, Jorge↗

Low-energy interband transition in the infrared response of the correlated metal SrVO 3 in the ultraclean limit

We studied the low-energy electronic response of the prototypical correlated metal SrVO 3 in the ultraclean and disordered limit using infrared spectroscopy and density functional theory plus dynamical mean field theory calculations (DFT+DMFT). A strong optical excitation at 70 meV is observed in the optical response of the ultraclean samples but is hidden by the low-energy Drude-like response from intraband excitations in the more disordered samples. DFT+DMFT calculations reveal that this optical excitation originates from interband transitions between the bands split by orbital off-diagonal hopping, which has often been ignored in cubic systems, such as SrVO 3 . A memory function analysis of the optical data shows that this interband transition can lead to deviations of optical self-energy from the expected Fermi-liquid behavior. Our findings demonstrate that analysis schemes employed to extract many-body effects from optical spectra may be oversimplified to study the true electronic ground state and that improvements in material quality can guide efforts to refine theoretical approaches.

36 MATERIALS SCIENCE↗

Enzyme Engineering Database (EnzEngDB): a platform for sharing and interpreting sequence–function relationships across protein engineering campaigns

The discovery and engineering of new enzymes is important across the bioeconomy, with diverse applications from foods to pharmaceuticals, sensors to agriculture. However, enzyme engineering, in particular machine learning-guided engineering, is hampered by a lack of data. Currently there exists no database designed to capture and interpret datasets created in this domain, nor are there easy analysis and visualisation tools. We developed the Enzyme Engineering Database to provide a centralized resource and an online analysis tool to consolidate sequence-function data from enzyme engineering campaigns, thereby making three contributions: (i) a database into which researchers can deposit public data, (ii) visualisation and analysis tools for protein engineers to analyse their own data or compare enzyme variants to other engineering campaigns, and (iii) a gold-standard dataset for benchmarking automated extraction along with the first large language model extraction pipeline specific for enzyme engineering campaigns. The Enzyme Engineering Database is accessible at http://enzengdb.org/.

Long, Yueming [California Institute of Technology ↗

Earth resources interactive processing system

System allows for processing and analysis of remotely-sensed Earth resources data. System may be modified for other sensors and allows numerous analysis functions on various types of image data.

Source record↗

GMI-IPS: Python Processing Software for Aircraft Campaigns

NASA's Atmospheric Tomography Mission (ATom) seeks to understand the impact of anthropogenic air pollution on gases in the Earth's atmosphere. Four flight campaigns are being deployed on a seasonal basis to establish a continuous global-scale data set intended to improve the representation of chemically reactive gases in global atmospheric chemistry models. The Global Modeling Initiative (GMI), is creating chemical transport simulations on a global scale for each of the ATom flight campaigns. To meet the computational demands required to translate the GMI simulation data to grids associated with the flights from the ATom campaigns, the GMI ICARTT Processing Software (GMI-IPS) has been developed and is providing key functionality for data processing and analysis in this ongoing effort. The GMI-IPS is written in Python and provides computational kernels for data interpolation and visualization tasks on GMI simulation data. A key feature of the GMI-IPS, is its ability to read ICARTT files, a text-based file format for airborne instrument data, and extract the required flight information that defines regional and temporal grid parameters associated with an ATom flight. Perhaps most importantly, the GMI-IPS creates ICARTT files containing GMI simulated data, which are used in collaboration with ATom instrument teams and other modeling groups. The initial main task of the GMI-IPS is to interpolate GMI model data to the finer temporal resolution (1-10 seconds) of a given flight. The model data includes basic fields such as temperature and pressure, but the main focus of this effort is to provide species concentrations of chemical gases for ATom flights. The software, which uses parallel computation techniques for data intensive tasks, linearly interpolates each of the model fields to the time resolution of the flight. The temporally interpolated data is then saved to disk, and is used to create additional derived quantities. In order to translate the GMI model data to the spatial grid of the flight path as defined by the pressure, latitude, and longitude points at each flight time record, a weighted average is then calculated from the nearest neighbors in two dimensions (latitude, longitude). Using SciPya's Regular Grid Interpolator, interpolation functions are generated for the GMI model grid and the calculated weighted averages. The flight path points are then extracted from the ATom ICARTT instrument file, and are sent to the multi-dimensional interpolating functions to generate GMI field quantities along the spatial path of the flight. The interpolated field quantities are then written to a ICARTT data file, which is stored for further manipulation. The GMI-IPS is aware of a generic ATom ICARTT header format, containing basic information for all flight campaigns. The GMI-IPS includes logic to edit metadata for the derived field quantities, as well as modify the generic header data such as processing dates and associated instrument files. The ICARTT interpolated data is then appended to the modified header data, and the ICARTT processing is complete for the given flight and ready for collaboration. The output ICARTT data adheres to the ICARTT file format standards V1.1. The visualization component of the GMI-IPS uses Matplotlib extensively and has several functions ranging in complexity. First, it creates a model background curtain for the flight (time versus model eta levels) with the interpolated flight data superimposed on the curtain. Secondly, it creates a time-series plot of the interpolated flight data. Lastly, the visualization component creates averaged 2D model slices (longitude versus latitude) with overlaid flight track circles at key pressure levels. The GMI-IPS consists of a handful of classes and supporting functionality that have been generalized to be compatible with any ICARTT file that adheres to the base class definition. The base class represents a generic ICARTT entry, only defining a single time entry and 3D spatial positioning parameters. Other classes inherit from this base class; several classes for input ICARTT instrument files, which contain the necessary flight positioning information as a basis for data processing, as well as other classes for output ICARTT files, which contain the interpolated model data. Utility classes provide functionality for routine procedures such as: comparing field names among ICARTT files, reading ICARTT entries from a data file and storing them in data structures, and returning a reduced spatial grid based on a collection of ICARTT entries. Although the GMI-IPS is compatible with GMI model data, it can be adapted with reasonable effort for any simulation that creates Hierarchical Data Format (HDF) files. The same can be said of its adaptability to ICARTT files outside of the context of the ATom mission. The GMI-IPS contains just under 30,000 lines of code, eight classes, and a dozen drivers and utility programs. It is maintained with GIT source code management and has been used to deliver processed GMI model data for the ATom campaigns that have taken place to date.

Damon, M. R.↗

Elastic depths for detecting shape anomalies in functional data

This this paper, we propose a new family of depth measures called the elastic depths that can be used to greatly improve shape anomaly detection in functional data. Shape anomalies are functions that have considerably different geometric forms or features from the rest of the data. Identifying them is generally more difficult than identifying magnitude anomalies because shape anomalies are often not distinguishable from the bulk of the data with visualization methods. The proposed elastic depths use the recently developed elastic distances to directly measure the centrality of functions in the amplitude and phase spaces. Measuring shape outlyingness in these spaces provides a rigorous quantification of shape which, in turn, gives the elastic depths a strong theoretical and practical advantage over other methods in detecting shape anomalies. A simple boxplot and thresholding method are introduced to identify shape anomalies using the elastic depths. We assess the elastic depth's detection skill on simulated shape outlier scenarios and compare them against popular shape anomaly detectors. Finally, bond yields, image outlines, and hurricane trajectories are used to demonstrate our method's applicability to functional data observed on three different manifolds.

97 MATHEMATICS AND COMPUTING↗

Trimming and Decontamination of Metagenomic Data can Significantly Impact Assembly and Binning Metrics, Phylogenomic and Functional Analysis

Background: Investigators using metagenomic sequencing to study microbiomes often trim and decontaminate reads without knowing their effect on downstream analyses. Objective: This study was designed to evaluate the impacts JGI trimming and decontamination procedures have on assembly and binning metrics, placement of MAGs into species trees, and functional profiles of MAGs extracted from complex rhizosphere metagenomes, as well as how more aggressive trimming impacts these binning metrics. Methods: Twenty-three Miscanthus x giganteus rhizosphere metagenomes were subjected to different combinations and thresholds of force, kmer, and quality trimming and decontamination using BBDuk. Reads were assembled and binned in KBase. Phylogenomic and statistical analyses were applied to evaluate the effects of trimming and decontamination on downstream analyses. Results: We found that JGI trimmed and decontaminated reads had significant impacts on assembly and binning metrics compared to raw reads, including significantly higher total contig counts, more contigs greater than 10k bp in length, and larger total lengths of raw assemblies compared to QC assemblies, and 2.0% lower average contamination of QC MAGs compared to raw MAGs. We also found that differences in the placement of MAGs in species trees increased with decreasing completeness and contamination thresholds. Furthermore, aggressive trimming (Q20) was found to significantly reduce MAG counts. Conclusion: Trimming and decontamination of metagenomics reads prior to assembly can change an investigator’s answer to the questions, “Who is there and what are they doing?” However, mild trimming and decontamination of metagenomic reads with high-quality scores are recommended for removing sample processing and sequencing artifacts.

Whitham, Jason M.↗

Likelihood Methods for CMB Experiments

A great deal of experimental effort is currently being devoted to the precise measurements of the cosmic microwave background (CMB) sky in temperature and polarization. Satellites, balloon-borne, and ground-based experiments scrutinize the CMB sky at multiple scales, and therefore enable to investigate not only the evolution of the early Universe, but also its late-time physics with unprecedented accuracy. The pipeline leading from time ordered data as collected by the instrument to the final product is highly structured. Moreover, it has also to provide accurate estimates of statistical and systematic uncertainties connected to the specific experiment. In this paper, we review likelihood approaches targeted to the analysis of the CMB signal at different scales, and to the estimation of key cosmological parameters. We consider methods that analyze the data in the spatial (i.e., pixel-based) or harmonic domain. We highlight the most relevant aspects of each approach and compare their performance.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Intraseasonal oscillations in the global atmosphere. I - Northern Hemisphere and tropics

Oscillatory modes in the Northern Hemisphere and in the tropics were examined systematically. The 700 mb heights were used to analyze extratropical oscillations, and the outgoing longwave radiation to study tropical oscillations in convection. All datasets were band-pass filtered to focus on the intraseasonal (IS) band of 10-120 days. Leading spatial patterns of variability were obtained by applying empirical orthogonal function analysis to these IS data. The leading principal components were subjected to singular spectrum analysis. In the Northern Hemisphere, there are two important modes of oscillation with periods near 48 and 23 days, respectively. The 48-day mode is the most important of the two. It has both traveling and standing components, and is dominated by a zonal wavenumber two. The 23-day mode has the spatial structure and propagation properties described by Branstator and by Kushnir (1987). In the tropics, the 40-50 day oscillation documented by Madden and Julian (1972), Weickmann (1983), Lau, and their colleagues, dominates the Indian and Pacific oceans from 60 deg E to the date line. From 170 deg W to 90 deg W, however, a 24-28 day oscillation is equally strong. The extratropical modes are often independent of, and sometimes lead, the tropical modes.

Ghil, Michael↗

The NASA JSC Hypervelocity Impact Test Facility (HIT-F)

The NASA Johnson Space Center Hypervelocity Impact Test Facility was created in 1980 to study the hypervelocity impact characteristics of composite materials. The facility consists of the Hypervelocity Impact Laboratory (HIRL) and the Hypervelocity Analysis Laboratory (HAL). The HIRL supports three different-size light-gas gun ranges which provide the capability of launching particle sizes from 100 micron spheres to 12.7 mm cylinders. The HAL performs three functions: (1) the analysis of data collected from shots in the HIRL, (2) numerical and analytical modeling to predict impact response beyond test conditions, and (3) risk and damage assessments for spacecraft exposed to the meteoroid and orbital debris environments.

Crews, Jeanne L.↗

Analysis of Suomi - NPP VIIRS Vignetting Functions Based on Yaw Maneuver Data

The Suomi NPP Visible Infrared Imager Radiometer Suite (VIIRS) reflective bands are calibrated on-orbit via reference to regular solar observations through a solar attenuation screen (SAS) and diffusely reflected off a Spectralon (Registered Trademark) panel. The degradation of the Spectralon panel BRDF due to UV exposure is tracked via a ratioing radiometer (SDSM) which compares near simultaneous observations of the panel with direct observations of the sun (through a separate attenuation screen). On-orbit, the vignetting functions of both attenuation screens are most easily measured when the satellite performs a series of yaw maneuvers over a short period of time (thereby covering the yearly angular variation of solar observations in a couple of days). Because the SAS is fixed, only the product of the screen transmission and the panel BRDF was measured. Moreover, this product was measured by both VIIRS detectors as well as the SDSM detectors (albeit at different reflectance angles off the Spectralon panel). The SDSM screen is also fixed; in this case, the screen transmission was measured directly. Corrections for instrument drift and degradation, solar geometry, and spectral effects were taken into consideration. The resulting vignetting functions were then compared to the pre-launch measurements as well as models based on screen geometry.

Xiong, Xiaoxiong↗

A Robotics Enabled Eddy Current Testing System for Autonomous Inspection of Heat Exchanger Tubes

The objective of the project is to develop a robotics enabled eddy current testing system (REECTS) in automatic probe deployment, inspection, and data acquisition and analysis. The main functions of the REECTS are to: 1) identify geometry and locations of heat exchange tubes with assistance of an imaging recognition system; 2) precisely control the position and motion speed of ECT probes by an adaptive control system; 3) facilitate data analysis and real-time decision making for autonomous inspection assisted by machine learning algorithms.

20 FOSSIL-FUELED POWER PLANTS↗

Subcentimeter-size particle distribution functions in planetary rings from Voyager radio and photopolarimeter occultation data

Analysis of measurements of the scattered and direct components of Voyager 1 radio occultation signals at 3.5 and 13 cm wavelengths yield estimates of the distribution functions of supracentimeter-size particles and thickness of relatively broad regions in Saturn's rings. If mearurements of signal amplitude at a shorter wavelength are combined with the previously analyzed data, the shape of the distribution functions characterizing the smaller particles can be constrained. If size distributions of arbitrary form were considered, many solutions are found that are consistent with the three available observations of signal amplitude. The best-fit power law was calculated to the three observations at three wavelengths for several of the embedded Saturn ringlets. Mie scattering theory predicts that the measured phase of the radio occultation signal is highly sensitive to particles ranging from 0.1 to 1.0 wavelengths in size, thus additional constraints on the subcentimeter-size distribution functions for both the Saturn and Uranus rings can in principle be derived from radio phase measurements.

Zebker, Howard A.↗

Guiding the choice of informatics software and tools for lipidomics research applications

Progress in mass spectrometry lipidomics has led to a rapid proliferation of studies across biology and biomedicine. These generate extremely large raw datasets requiring sophisticated solutions to support automated data processing. To address this, numerous software tools have been developed and tailored for specific tasks. However, for researchers, deciding which approach best suits their application relies on ad hoc testing, which is inefficient and time consuming. Here we first review the data processing pipeline, summarizing the scope of available tools. Next, to support researchers, LIPID MAPS provides an interactive online portal listing open-access tools with a graphical user interface. This guides users towards appropriate solutions within major areas in data processing, including (1) lipid-oriented databases, (2) mass spectrometry data repositories, (3) analysis of targeted lipidomics datasets, (4) lipid identification and (5) quantification from untargeted lipidomics datasets, (6) statistical analysis and visualization, and (7) data integration solutions. Detailed descriptions of functions and requirements are provided to guide customized data analysis workflows.

59 BASIC BIOLOGICAL SCIENCES↗