Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “online algorithm”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

462 records · Page 26

Application of ML/AI for Identifying Earth Science Datasets in Research Publications

NASA Data Active Archive Centers, or DAACs, ingest, store and distribute data acquired from satellites, ground systems as well as modelling data. These data are organized by the datasets, each presenting collection of files usually associated with the certain mission, instrument, processing level, parameter(s), algorithm and/or model. The number of datasets offered by a single DAAC to the public varies. GES DISC, for example, currently offers for public use approximately ~1,300 datasets. While each publicly offered dataset comes with supporting documentation, it is challenging for novice and even experienced scientists to navigate among the datasets that offer similar parameters to find the datasets for their particular research application. Supplying dataset documentation with the scientific paper citations that refer to that dataset provides means for the dataset users to educate themselves with the application research that dataset is being used in. Collecting citations of the papers that use the datasets for their research yield valuable insights into application areas of those datasets, information about usage of the dataset groups for specific applications and those application topics. It also gives insights into the “deep metrics” of the dataset usage, as opposed to the common metrics of the dataset usage such as number of users who downloaded the dataset files and volumes of downloaded data. Association of a certain scientific paper with the dataset(s) presents a challenge because most of the paper authors do not properly cite the datasets, datasets usually have cryptic names and Digital Object Identifiers (DOIs) that are used for dataset identification were assigned to the datasets only few years ago. Simple Google or online library search do not provide even meaningful fraction of the results when performed by the dataset name or DOI, however they provide too many results when the search is done by more broader terms such as mission and instrument names. Attempts to create an AI system capable to identify dataset in the scientific papers have already been made using neural networks classifiers on the basis of the dataset mission, instrument and variable name. This method was applied to NASA SEDAC, which has 41 datasets in total. In GES DISC there can be as many as ~100 datasets per mission/instrument with some of the datasets consisting of multiple variables so there is a need for more differentiating parameters for dataset identification in the paper. The approach we are currently investigating is creating AI classifiers that are based on multiple dataset features, or keywords, extracted from the NASA Earthdata Common Dataset Repository (CMR). The features are weighted based on how precisely they can identify a dataset. The classifier uses preprocessed paper text as input and searches for the CMR datasets whose feature sets are the closest to the feature sets contained in the paper. The challenges of dataset identification include variety of ways the paper authors describe the datasets in their papers and incomplete tagging of the CMR dataset description (DIFs).

Irina Gerasimov↗

Classifying Seyfert Galaxies with Deep Learning

The traditional classification for a subclass of the Seyfert galaxies is visual inspection or using a quantity defined as a flux ratio between the Balmer line and forbidden line. One algorithm of deep learning is the convolution neural network (CNN), which has shown successful classification results. We build a one-dimensional CNN model to distinguish Seyfert 1.9 spectra from Seyfert 2 galaxies. We find that our model can recognize Seyfert 1.9 and Seyfert 2 spectra with an accuracy of over 80% and pick out an additional Seyfert 1.9 sample that was missed by visual inspection. We use the new Seyfert 1.9 sample to improve the performance of our model and obtain a 91% precision of Seyfert 1.9. These results indicate that our model can pick out Seyfert 1.9 spectra among Seyfert 2 spectra. We decompose the Hα emission line of our Seyfert 1.9 galaxies by fitting two Gaussian components and derive the line width and flux. We find that the velocity distribution of the broad Hα component of the new Seyfert 1.9 sample has an extending tail toward the higher end, and the luminosity of the new Seyfert 1.9 sample is slightly weaker than the original Seyfert 1.9 sample. This result indicates that our model can pick out the sources that have a relatively weak broad Hα component. In addition, we check the distributions of the host galaxy morphology of our Seyfert 1.9 samples and find that the distribution of the host galaxy morphology is dominated by a large bulge galaxy. In the end, we present an online catalog of 1297 Seyfert 1.9 galaxies with measurements of the Hα emission line.

79 ASTRONOMY AND ASTROPHYSICS↗

Serious Gaming for Building a Basis of Certification via Trust and Trustworthiness of Autonomous Systems

Autonomous systems governed by a variety of adaptive and nondeterministic algorithms are being planned for inclusion into safety-critical environments, such as unmanned aircraft and space systems in both civilian and military applications. However, until autonomous systems are proven and perceived to be capable and resilient in the face of unanticipated conditions, humans will be reluctant or unable to delegate authority, remaining in control aided by machine-based information and decision support. Proving capability, or trustworthiness, is a necessary component of certification. Perceived capability is a component of trust. Trustworthiness is an attribute of a cyber-physical system that requires context-driven metrics to prove and certify. Trust is an attribute of the agents participating in the system and is gained over time and multiple interactions through trustworthy behavior and transparency. Historically, artificial intelligence and machine learning systems provide answers without explanation - without a rationale or insight into the machine “thinking”. In order to function as trusted teammates, machines must be able to explain their decisions and actions. This transparency is a product of both content and communication. NASA’s Autonomy Teaming & TRAjectories for Complex Trusted Operational Reliability (ATTRACTOR) project seeks to build a basis for certification of autonomous systems via establishing metrics for trustworthiness and trust in multi-agent team interactions, using AI (Artificial Intelligence) explainability and persistent modeling and simulation, in the context of mission planning and execution, with analyzable trajectories. Inspired by Massively Multiplayer Online Role Playing Games (MMORPG) and Serious Gaming, the proposed ATTRACTOR modeling and simulation environment is similar to online gaming environments in which player (aka agent) participants interact with each other, affect their environment, and expect the simulation to persist and change regardless of any individual agent’s active participation. This persistent simulation environment will accommodate individual agents, groups of self-organizing agents, and large-scale infrastructure behavior. The effects of the emerging adaptation and coevolution can be observed and measured to building a basis of measurable trustworthiness and trust, toward certification of safety-critical autonomous systems.

Allen, B. Danette↗

Online model-based diagnosis to support autonomous operation of an advanced life support system

This article describes methods for online model-based diagnosis of subsystems of the advanced life support system (ALS). The diagnosis methodology is tailored to detect, isolate, and identify faults in components of the system quickly so that fault-adaptive control techniques can be applied to maintain system operation without interruption. We describe the components of our hybrid modeling scheme and the diagnosis methodology, and then demonstrate the effectiveness of this methodology by building a detailed model of the reverse osmosis (RO) system of the water recovery system (WRS) of the ALS. This model is validated with real data collected from an experimental testbed at NASA JSC. A number of diagnosis experiments run on simulated faulty data are presented and the results are discussed.

Non-NASA Center↗

Dictionary Learning with Accumulator Neurons

The Locally Competitive Algorithm (LCA) uses local competition between non-spiking leaky integrator neurons to infer sparse representations, allowing for potentially real-time execution on massively parallel neuromorphic architectures such as Intel's Loihi processor. Here, we focus on the problem of inferring sparse representations from streaming video using dictionaries of spatiotemporal features optimized in an unsupervised manner for sparse reconstruction. Non-spiking LCA has previously been used to achieve unsupervised learning of spatiotemporal dictionaries composed of convolutional kernels from raw, unlabeled video. We demonstrate how unsupervised dictionary learning with spiking LCA (\hbox{S-LCA}) can be efficiently implemented using accumulator neurons, which combine a conventional leaky-integrate-and-fire (\hbox{LIF}) spike generator with an additional state variable that is used to minimize the difference between the integrated input and the spiking output. We demonstrate dictionary learning across a wide range of dynamical regimes, from graded to intermittent spiking, for inferring sparse representations of both static images drawn from the CIFAR database as well as video frames captured from a DVS camera. On a classification task that requires identification of the suite from a deck of cards being rapidly flipped through as viewed by a DVS camera, we find essentially no degradation in performance as the LCA model used to infer sparse spatiotemporal representations migrates from graded to spiking. We conclude that accumulator neurons are likely to provide a powerful enabling component of future neuromorphic hardware for implementing online unsupervised learning of spatiotemporal dictionaries optimized for sparse reconstruction of streaming video from event based DVS cameras.

artificial intelligence↗

A Satellite-Derived Climate-Quality Data Record of the Clear-Sky Surface Temperature of the Greenland Ice Sheet

We have developed a climate-quality data record of the clear-sky surface temperature of the Greenland Ice Sheet using the Moderate-Resolution Imaging Spectroradiometer (MODIS) Terra ice-surface temperature (1ST) algorithm. A climate-data record (CDR) is a time series of measurements of sufficient length, consistency, and continuity to determine climate variability and change. We present daily and monthly Terra MODIS ISTs of the Greenland Ice Sheet beginning on 1 March 2000 and continuing through 31 December 2010 at 6.25-km spatial resolution on a polar stereographic grid within +/-3 hours of 17:00Z or 2:00 PM Local Solar Time. Preliminary validation of the ISTs at Summit Camp, Greenland, during the 2008-09 winter, shows that there is a cold bias using the MODIS IST which underestimates the measured surface temperature by approximately 3 C when temperatures range from approximately -50 C to approximately -35 C. The ultimate goal is to develop a CDR that starts in 1981 with the Advanced Very High Resolution (AVHRR) Polar Pathfinder (APP) dataset and continues with MODIS data from 2000 to the present. Differences in the APP and MODIS cloud masks have so far precluded the current IST records from spanning both the APP and MODIS IST time series in a seamless manner though this will be revisited when the APP dataset has been reprocessed. The Greenland IST climate-quality data record is suitable for continuation using future Visible Infrared Imager Radiometer Suite (VIIRS) data and will be elevated in status to a CDR when at least 9 more years of climate-quality data become available either from MODIS Terra or Aqua, or from the VIIRS. The complete MODIS IST data record will be available online in the summer of 2011.

Hall, Dorothy K.↗

Effects of overlapping sources on cosmic shear estimation: Statistical sensitivity and pixel-noise bias

The next generation of dark-energy imaging surveys — so called “Stage-IV” surveys, such as that of the Rubin Observatory Legacy Survey of Space and Time (LSST) — will cross a threshold in the number density of detected sources on the sky that requires qualitatively different image analysis and measurement techniques compared to the current generation of Stage-III surveys. In Stage-IV surveys, a significant amount of the cosmologically useful information is due to sources whose images overlap with those of other sources on the sky. Here, we focus on the weak gravitational lensing probe, for which we expect the largest impact since the cosmic shear signal is primarily encoded in the estimated shapes of observed galaxies and thus directly impacted by overlaps. We introduce a framework based on the Fisher formalism to analyze the effect of the overlapping sources (“blending”) on the estimation of cosmic shear. This method gives concrete predictions for the minimum loss of information due to noise and blending for any choice of “deblending” scheme and shape-measurement algorithm. Our studies account for undetected sources but do not address their full effects and biases they may introduce. We use simulated images and predict this impact of blending for three surveys: the Dark Energy Survey (DES), the Hyper-Suprime Cam Subaru Strategic Program (HSC-SSP), and the Rubin LSST. Our methodology successfully estimates the statistical sensitivity to weak lensing for DES and HSC early results. For LSST, we present the expected loss in statistical sensitivity for the ten-year survey due to blending. We find that for approximately 62% of galaxies that are likely to be detected in full-depth LSST images, at least 1% of the flux in their pixels is from overlapping sources. We also find that the statistical correlations between measures of overlapping galaxies and, to a much lesser extent (0.2%) the higher shot noise level due to their presence, decrease the effective number density of galaxies, N eff , by ~ 18%. We calculate an upper limit on N eff of 39.4 galaxies per arcmin 2 in r band. We study the impact of stars on as a function of stellar density and illustrate the diminishing returns of extending the survey into lower Galactic latitudes. We extend the simulation-based Fisher formalism to predict the expected increase in pixel-noise bias due to blending for maximum-likelihood (ML) shape estimators. We find that noise bias depends sensitively on the particular shape estimator and measure of ensemble-average shape that is used, and properties of the galaxy that include redshift-dependent quantities such as size and luminosity. The source code for these studies is available online.[The documented software developed for the catalog-level studies are available in the open-source LSST DESC github repository https://github.com/LSSTDESC/WeakLensingDeblending. The software for analyzing one or two galaxies with user-defined parameters is in the open-source github repository https://github.com/ismael-mendoza/ShapeMeasurementFisherFormalism.]

79 ASTRONOMY AND ASTROPHYSICS↗

Data-driven linear time advance operators for the acceleration of plasma physics simulation

In this study, we demonstrate the application of data-driven linear operator construction for time advance with a goal of accelerating plasma physics simulation. We apply dynamic mode decomposition (DMD) to data produced by the nonlinear SOLPS-ITER (Scrape-off Layer Plasma Simulator - International Thermonuclear Experimental Reactor) plasma boundary code suite in order to estimate a series of linear operators and monitor their predictive accuracy via online error analysis. We find that this approach defines when these dynamics can be represented by a sequence of approximate linear operators and is essential for providing consistent projections when compared to an unconstrained application. For linear diffusion and advection–diffusion fluid test problems, we construct and apply operators within explicit and implicit time advance schemes, demonstrating that stability can be robustly guaranteed in each case. We further investigate the use of the linear time advance operators within several integration methods including forward Euler, backward Euler, and the matrix exponential. The application of this method to simulation data from SOLPS-ITER, with varying levels of Markov chain Monte Carlo numerical noise, shows that constrained DMD operators yield a capability to identify, extract, and integrate a (slow) subset of the present timescales. Example applications show that for projected speedup factors of [Formula: see text], and [Formula: see text], a mean relative error of 3%, 5%, and 8% and maximum relative error less than 20% are achievable, which appears acceptable for typical SOLPS-ITER steady-state simulations.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Initial Laboratory Demonstration of Multi-Star Wavefront Control at the Occulting Mask Coronagraph Testbed

Online Abstract: A majority of Sun-like stars, such as the A and B components of Alpha Centauri, have at least one stellar companion that can introduce additional noise into the field of view of any high-contrast imaging instrument. Multi-Star Wavefront Control (MSWC) is a wavefront-control technique that removes stellar leakage from both stellar components, enabling direct imaging of planets in many binary star systems. We present the latest experimental and modeling results obtained with MSWC as part of the technology development effort focusing on demonstrations conducted on the Occulting Mask Coronagraph (OMC) testbed at JPL. OMC has a layout similar to the Roman Space Telescope coronagraph instrument (CGI), and we used a MSWC mask similar to the one that was contributed to the Roman CGI. Our results represent the first demonstrations of this technique on the OMC testbed, with the ultimate goal of demonstrating full MSWC with validated models at contrast levels relevant to Roman CGI. Technical Review Abstract: A majority of Sun-like stars have at least one stellar companion that can introduce additional noise into the field of view of any high-contrast imaging instrument, limiting the achievable contrast. These include high-quality target stars such as the A and B components of Alpha Centauri, our nearest stellar neighbor. Enabling direct imaging of binary stars has the potential to increase the scientific yield for coronagraphic instruments planned on NASA's future space missions including the Roman Space Telescope and the next IR/O/UV Flagship recommended by Astro2020. Multi-Star Wavefront Control (MSWC) is a wavefront-control technique that simultaneously removes the (mutually incoherent) stellar leakage from both stellar components, enabling direct imaging of planets in many binary star systems. MSWC is an algorithmic technique and can be used with existing wavefront control systems on coronagraphic instruments (as well as starshades if a deformable mirror is available in the optical path). We summarize the latest experimental and numerical results obtained with MSWC as part of the technology development effort to demonstrate compatibility with existing high-contrast imaging platforms for this technique. The Super-Nyquist regime of MSWC was tested in vacuum at JPL's High Contrast Imaging Testbed (HCIT) on the Decadal Survey Testbed (DST) reaching 8.6e-9 contrast in a 10% band. Recently, the Occulting Mask Coronagraph (OMC) testbed at JPL is being prepared for demonstrations of Multi-Star Wavefront Control. A shaped pupil mask similar to the contributed MSWC mask on the Roman Space Telescope's coronagraph instrument has been recently manufactured including matching Lyot and focal plane masks and being installed on the OMC testbed. The goal of this experiment is a demonstration of MSWC using validated models on a testbed configuration and at contrast levels relevant to the Roman coronagraphic instrument.

High-contrast imaging↗

Strawman Philosophical Guide for Developing International Network of GPM GV Sites

The creation of an international network of ground validation (GV) sites that will support the Global Precipitation Measurement (GPM) Mission's international science programme will require detailed planning of mechanisms for exchanging technical information, GV data products, and scientific results. An important component of the planning will be the philosophical guide under which the network will grow and emerge as a successful element of the GPM Mission. This philosophical guide should be able to serve the mission in developing scientific pathways for ground validation research which will ensure the highest possible quality measurement record of global precipitation products. The philosophical issues, in this regard, partly stem from the financial architecture under which the GV network will be developed, i.e., each participating country will provide its own financial support through committed institutions -- regardless of whether a national or international space agency is involved.At the 1st International GPM Ground Validation Workshop held in Abingdon, UK in November-2003, most of the basic tenants behind the development of the international GV network were identified and discussed. Therefore, with this progress in mind, this presentation is intended to put forth a strawman philosophical guide supporting the development of the international network of GPM GV sites, noting that the initial progress has been reported in the Proceedings of the 1st International GPM GV Workshop -- available online. The central philosophical issues themselves, all flow from the fact that each participating institution can only bring to the table, GV facilities and scientific personnel that are affordable to the sanctioning (funding) national agency (be that a research, research-support, or operational agency). This situation imposes on the network, heterogeneity in the measuring sensors, data collection periods, data collection procedures, data latencies, and data reporting capabilities. Therefore, in order for the network to be effective in supporting the central scientific goals of the GPM mission, there must be a basic agreed upon doctrine under which the network participants function vis-a-vis: (1) an overriding set of general scientific requirements, (2) a minimal set of policies governing the free flow of GV data between the scientific participants, (3) a few basic definitions concerning the prioritization of measurements and their respective value to the mission, (4) a few basic procedures concerning data formats, data reporting procedures, data access, and data archiving, and (5) a simple means to differentiate GV sites according to their level of effort and ability to perform near real-time data acquisition - data reporting tasks. Most important, in case they choose to operate as a near real-time data collection-data distribution site, they would be expected to operate under a fairly narrowly defined protocol needed to ensure smooth GV support operations. This presentation will suggest measures responsive to items (1) - (5) from which to proceed,. In addition, this presentation will seek to stimulate discussion and debate concerning how much heterogeneity is tolerable within the eventual GV site network, given that the any individual GV site can only be considered scientifically useful if it supports the achievement of the central GPM Mission goals. Only ground validation research that has a direct connection to the space mission should be considered justifiable given the overarching scientific goals of the mission. Therefore each site will have to seek some level of accommodation to what the GPM Mission requires in the way of retrieval error characterization, retrieval error detection and reporting, and generation of GV data products that support assessment and improvement of the mission's standard precipitation retrieval algorithms. These are all important scientific issues that will be best resolved in open scientific debate.

Smith, Eric A.↗

Track reconstruction as a service for collider physics

Optimizing charged-particle track reconstruction algorithms is crucial for efficient event reconstruction in Large Hadron Collider (LHC) experiments due to their significant computational demands. Existing track reconstruction algorithms have been adapted to run on massively parallel coprocessors, such as graphics processing units (GPUs), to reduce processing time. Nevertheless, challenges remain in fully harnessing the computational capacity of coprocessors in a scalable and non-disruptive manner. This paper proposes an inference-as-a-service approach for particle tracking in high energy physics experiments. To evaluate the efficacy of this approach, two distinct tracking algorithms are tested: Patatrack, a rule-based algorithm, and Exa.TrkX, a machine learning-based algorithm. The as-a-service implementations show enhanced GPU utilization and can process requests from multiple CPU cores concurrently without increasing per-request latency. The impact of data transfer is minimal and insignificant compared to running on local coprocessors. This approach greatly improves the computational efficiency of charged particle tracking, providing a solution to the computing challenges anticipated in the High-Luminosity LHC era.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Phasor-Measurement-Unit-Based Data Analytics Using Digital Twin and PhasorAnalytics Software

A major objective of this project was to apply GE’s commercial machine learning and data analytics toolsets to large-scale, real-world, anonymized Phasor Measurement Unit (PMU) datasets in order to extract signatures, correlated and/or causal factors, and precursor patterns associated with significant power system phenomena. The project had a particular emphasis on extraction of insights relevant to asset health monitoring, real-time load modeling and cybersecurity monitoring. Additionally, the team was directed to undertake a comprehensive data quality analysis for the provided datasets and encouraged to estimate the ‘machine-learning readiness’ of the datasets by documenting any major obstacles to the application of commercial machine learning algorithms. To accomplish the aforementioned objectives, the project team’s work centered around the identification of key event signatures and application of the identified event signatures for event detection and event classification. The industry-validated, semi-supervised machine learning strategy employed for event signature identification involved several major tasks, including data-preprocessing, generation of an overabundance of features, normal data identification, normality modeling, and event signature identification through a methodical, quantitative ranking of features in order of relevance to each studied event type. Throughout the project, data quality issues and mitigation techniques were investigated. In this report, insights are provided regarding the readiness of the provided synchrophasor datasets for application of machine learning and data analytics. The methodologies employed for this technical strategy are summarized in this report. With regards to data preprocessing and feature generation, the provided Training and Test Datasets were ingested into GE’s big data environment. Subsequently, the team applied bad data cleansing and data imputation scripts, event detection scripts, and application programming interfaces (APIs) to the datasets for convenient data access. The project team completed development and validation of dozens of physics-based, statistics-based and transformation-based feature functions used for the extraction of over 60 synchrophasor features. Using a new parallel feature generation technology developed on this project, over 60 features have been rapidly generated for the full two years’ worth of Training and Test Dataset data associated with both the Eastern and Western interconnects. Even accommodating for temporal down-sampling inherent to the feature extraction procedure, this parallel feature generation activity resulted in a massive feature set with a storage requirement approximately equal to that of the raw training dataset itself. With regards to normal data identification and normality modeling, a normality model was built using the feature data extracted from the Training Dataset and iteratively refined subsequent to incremental adjustments and expansions of the Training Dataset feature data. With respect to event characterization and signature identification, an event signature identification pipeline was developed and used in conjunction with the normality model to identify over 15 event signatures for key event categories within the Training Dataset. The identified event signatures were used to characterize hundreds of key events in terms of relative severity, duration, and location of the event. An investigation was undertaken to identify correlated and causal factors involved in transformer events. A separate investigation into temporal trends in ring-down analysis results was undertaken to determine possible associations between system dynamics and various other factors such as loading, season or year. To validate the identified event signatures, additional work was undertaken to develop signature-based anomaly detection and classification tools suitable for convenient application to the synchrophasor datasets. The anomaly detection and classification tools, suitable for online application, were then applied to the entirety of the Eastern Interconnect Training and Test Datasets. Performance of the event detection and classification tools was evaluated upon receipt of the Test Dataset event logs (i.e., the labels for events contained in the Test Dataset), and promising results were obtained despite several challenges (documented herein) associated with application of supervised or semi-supervised machine learning methods to large-scale, anonymized datasets. Finally, the detection and classification tools were used to detect, classify, and characterize thousands of new events not included in the original event logs provided by the DOE within both the Training and Test Datasets.

24 POWER TRANSMISSION AND DISTRIBUTION↗