Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Data mining”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

Asi Nuclear Energy Sensors Data Portal Chatbot And Data Structuring Tool

The Idaho National Laboratory (INL) is advancing the development of an AI-powered chatbot and data structuring tool specifically designed to accelerate data mining processes for sensor-related information and seamlessly integrate the results into the ASI Sensors Data Portal (https://nes.energy.gov/). By doing so, the software aims to enhance the accessibility, usability, and organization of sensor data for nuclear energy applications. The software initial phase focuses on retrieving comprehensive datasets, prioritizing the past five years of publicly available information from the Office of Scientific and Technical Information (OSTI). These datasets will be meticulously processed to ensure compatibility, employing cleaning and preprocessing steps to eliminate irrelevant, incomplete, or corrupted information, thus establishing a robust foundation for subsequent AI use. The data will serve as the backbone for training an AI model and chatbot, which will act as an interactive tool enabling users to ask complex, context-specific questions and receive accurate, validated answers derived from constrained literature. In parallel, the project incorporates a data structuring process supported by AI to organize sensor information from multiple sources into a standardized format. This structured data will include detailed sensor specifications, such as measurement range, applications, accuracy, and operating conditions, generated and documented with AI. These specifications will be systematically integrated into the sensor portal. To maintain the highest levels of accuracy and relevance, all AI-generated outputs will be reviewed and validated by subject matter experts (SMEs), with additional fields or parameters added as needed. Future stages of the project aim to expand the dataset beyond OSTI to include other sources and potentially incorporate unclassified controlled information (UCI) with restricted access protocols to address security and confidentiality requirements.

Mapes, NormanJ. [Idaho National Laboratory (INL), ↗

Compiler and Runtime Approaches to Enable Large-Scale Irregular Programs. Final report, July 2013 - July 2019

While regular algorithms, characterized by operations on dense matrices and arrays, have long been the mainstay of scientific, high-performance computing, irregular algorithms, which feature unpredictable accesses to pointer-based data structures, are becoming increasingly common in high performance computing, arising in graph analysis, data mining and visualization, among other domains. Unfortunately, the defining characteristics of irregular applications, their dynamic, unpredictable, data-dependent access patterns and data layouts, make achieving high performance on large scale systems difficult. Scaling applications to peta- and exa-scale requires carefully controlling communication and data movement and placement, an inherently difficult task when access patterns and data layouts are unpredictable! Most irregular applications that attain high performance must be painstakingly hand-written and hand-tuned, with few common principles or paradigms uniting various implementations and easing future development. Despite the increasing importance of irregular applications, there is little programmer knowledge, and even less compiler ability, devoted to optimizing them. This project aims to solve these problems. By allowing programmers to write irregular applications in high level forms, with at most a few annotations highlighting key structural properties, programmers can focus on developing their algorithms and methods. The compiler and run-time system can take on the tedious task of optimizing the application for execution at large scales, and can automatically provide efficient implementations. This will provide portability and ease maintenance for existing irregular applications, but, more importantly, open up whole new domains of computational science to large-scale, high-performance simulation codes.

97 MATHEMATICS AND COMPUTING↗

Objective Assessment Method for RNAV STAR Adherence

Flight crews and air traffic controllers have reported many safety concerns regarding area navigation standard terminal arrival routes (RNAV STARs). Specifically, optimized profile descents (OPDs). However, our information sources to quantify these issues are limited to subjective reporting and time consuming case-by-case investigations. This work is a preliminary study into the objective performance of instrument procedures and provides a framework to track procedural concepts and assess design specifications. We created a tool and analysis methods for gauging aircraft adherence as it relates to RNAV STARs. This information is vital for comprehensive understanding of how our air traffic behaves. In this study, we mined the performance of 24 major US airports over the preceding three years. Overlaying 4D radar track data onto RNAV STAR routes provided a comparison between aircraft flight paths and the waypoint positions and altitude restrictions. NASA Ames Supercomputing resources were utilized to perform the data mining and processing. We assessed STARs by lateral transition path (full-lateral), vertical restrictions (full-lateral/full-vertical), and skipped waypoints (skips). In addition, we graphed frequencies of aircraft altitudes relative to the altitude restrictions. Full-lateral adherence was always greater than Full-lateral/ full- vertical, as it is a subset, but the difference between the rates was not consistent. Full-lateral/full-vertical adherence medians of the 2016 procedures ranged from 0% in KDEN (Denver) to 21% in KMEM (Memphis). Waypoint skips ranged from 0% to nearly 100% for specific waypoints. Altitudes restrictions were sometimes missed by systematic amounts in 1,000 ft. increments from the restriction, creating multi-modal distributions. Other times, altitude misses looked to be more normally distributed around the restriction. This tool may aid in providing acceptability metrics as well as risk assessment information.

Stewart, Michael↗

A chemistry-informed hybrid machine learning approach to predict metal adsorption onto mineral surfaces

Historically, surface complexation model (SCM) constants and distribution coefficients (K d ) have been employed to quantify mineral-based retardation effects controlling the fate of metals in subsurface geologic systems. Our recent SCM development workflow, based on the Lawrence Livermore National Laboratory Surface Complexation/Ion Exchange (L-SCIE) database, illustrated a community FAIR data approach to SCM development by predicting uranium(VI)-quartz adsorption for a large number of literature-mined data. Here, we present an alternative hybrid machine learning (ML) approach that shows promise in achieving equivalent high-quality predictions compared to traditional surface complexation models. At its core, the hybrid random forest (RF) ML approach is motivated by the proliferation of incongruent SCMs in the literature that limit their applicability in reactive transport models. Our hybrid ML approach implements PHREEQC-based aqueous speciation calculations; values from these simulations are automatically used as input features for a random forest (RF) algorithm to quantify adsorption and avoid SCM modeling constraints entirely. Named the LLNL Speciation Updated Random Forest (L-SURF) model, this hybrid approach is shown to have applicability to U(VI) sorption cases driven by both ion-exchange and surface complexation, as is shown for quartz and montmorillonite cases. The approach can be applied to reactive transport modeling and may provide an alternative to the costly development of self-consistent SCM reaction databases.

38 RADIATION CHEMISTRY, RADIOCHEMISTRY, AND NUCLEA↗

A Sparse Tensor Benchmark Suite for CPUs and GPUs

Tensor computations present significant performance chal- lenges that impact a wide spectrum of applications ranging from machine learning, healthcare analytics, social network analysis, data mining to quantum chemistry and signal processing. Efforts to improve the perfor- mance of tensor computations include exploring data layout, execution scheduling, and parallelism in common tensor kernels. This work presents a benchmark suite for arbitrary-order sparse tensor kernels using state- of-the-art tensor formats: coordinate (COO) and hierarchical coordinate (HiCOO) on CPUs and GPUs. It presents a set of reference tensor kernel implementations that are compatible with real-world tensors and power law tensors extended from synthetic graph generation techniques. We also propose Roofline performance models for these kernels to provide insights of computer platforms from sparse tensor view. This benchmark suite along with the synthetic tensor generator is publicly available.

Li, Jiajia↗

Access to Archived Astronaut Data for Human Research Program Researchers: Update on Progress and Process Improvements

Since the 2010 NASA directive to make the Life Sciences Data Archive (LSDA) and Lifetime Surveillance of Astronaut Health (LSAH) data archives more accessible by the research and operational communities, demand for astronaut medical data has increased greatly. LSAH and LSDA personnel are working with Human Research Program on many fronts to improve data access and decrease lead time for release of data. Some examples include the following: Feasibility reviews for NASA Research Announcement (NRA) data mining proposals; Improved communication, support for researchers, and process improvements for retrospective Institutional Review Board (IRB) protocols; Supplemental data sharing for flight investigators versus purely retrospective studies; Work with the Multilateral Human Research Panel for Exploration (MHRPE) to develop acceptable data sharing and crew consent processes and to organize inter-agency data coordinators to facilitate requests for international crewmember data. Current metrics on data requests crew consenting will be presented, along with limitations on contacting crew to obtain consent. Categories of medical monitoring data available for request will be presented as well as flow diagrams detailing data request processing and approval steps.

Lee, L. R.↗

Text-mined dataset of gold nanoparticle synthesis procedures, morphologies, and size entities

Abstract Gold nanoparticles are highly desired for a range of technological applications due to their tunable properties, which are dictated by the size and shape of the constituent particles. Many heuristic methods for controlling the morphological characteristics of gold nanoparticles are well known. However, the underlying mechanisms controlling their size and shape remain poorly understood, partly due to the immense range of possible combinations of synthesis parameters. Data-driven methods can offer insight to help guide understanding of these underlying mechanisms, so long as sufficient synthesis data are available. To facilitate data mining in this direction, we have constructed and made publicly available a dataset of codified gold nanoparticle synthesis protocols and outcomes extracted directly from the nanoparticle materials science literature using natural language processing and text-mining techniques. This dataset contains 5,154 data records, each representing a single gold nanoparticle synthesis article, filtered from a database of 4,973,165 publications. Each record contains codified synthesis protocols and extracted morphological information from a total of 7,608 experimental and 12,519 characterization paragraphs.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Multi-kernel Edge Attention Graph Autoencoder

MEAGraph (Multi-kernel Edge Attention Graph Autoencoder) is a graph-based autoencoder model designed for unsupervised data mining for datasets used in machine learning potentials. It provides accurate clustering for atomic environment identification, unsupervised and unlabeled data pruning for dataset construction.

Sun, Hong↗

Autonomous, Context-Sensitive, Task Management Systems and Decision Support Tools I: Human-Autonomy Teaming Fundamentals and State of the Art

Recent advances in artificial intelligence, machine learning, data mining and extraction, and especially in sensor technology have resulted in the availability of a vast amount of digital data and information and the development of advanced automated reasoners. This creates the opportunity for the development of a robust dynamic task manager and decision support tool that is context sensitive and integrates information from a wide array of on-board and off aircraft sourcesa tool that monitors systems and the overall flight situation, anticipates information needs, prioritizes tasks appropriately, keeps pilots well informed, and is nimble and able to adapt to changing circumstances. This is the first of two companion reports exploring issues associated with autonomous, context-sensitive, task management and decision support tools. In the first report, we explore fundamental issues associated with the development of an integrated, dynamic, flight information and automation management system. We discuss human factors issues pertaining to information automation and review the current state of the art of pilot information management and decision support tools. We also explore how effective human-human team behavior and expectations could be extended to teams involving humans and automation or autonomous systems.

context-sensitive↗

Spinoff 2013

Topics covered include: Innovative Software Tools Measure Behavioral Alertness; Miniaturized, Portable Sensors Monitor Metabolic Health; Patient Simulators Train Emergency Caregivers; Solar Refrigerators Store Life-Saving Vaccines; Monitors Enable Medication Management in Patients' Homes; Handheld Diagnostic Device Delivers Quick Medical Readings; Experiments Result in Safer, Spin-Resistant Aircraft; Interfaces Visualize Data for Airline Safety, Efficiency; Data Mining Tools Make Flights Safer, More Efficient; NASA Standards Inform Comfortable Car Seats; Heat Shield Paves the Way for Commercial Space; Air Systems Provide Life Support to Miners; Coatings Preserve Metal, Stone, Tile, and Concrete; Robots Spur Software That Lends a Hand; Cloud-Based Data Sharing Connects Emergency Managers; Catalytic Converters Maintain Air Quality in Mines; NASA-Enhanced Water Bottles Filter Water on the Go; Brainwave Monitoring Software Improves Distracted Minds; Thermal Materials Protect Priceless, Personal Keepsakes; Home Air Purifiers Eradicate Harmful Pathogens; Thermal Materials Drive Professional Apparel Line; Radiant Barriers Save Energy in Buildings; Open Source Initiative Powers Real-Time Data Streams; Shuttle Engine Designs Revolutionize Solar Power; Procedure-Authoring Tool Improves Safety on Oil Rigs; Satellite Data Aid Monitoring of Nation's Forests; Mars Technologies Spawn Durable Wind Turbines; Programs Visualize Earth and Space for Interactive Education; Processor Units Reduce Satellite Construction Costs; Software Accelerates Computing Time for Complex Math; Simulation Tools Prevent Signal Interference on Spacecraft; Software Simplifies the Sharing of Numerical Models; Virtual Machine Language Controls Remote Devices; Micro-Accelerometers Monitor Equipment Health; Reactors Save Energy, Costs for Hydrogen Production; Cameras Monitor Spacecraft Integrity to Prevent Failures; Testing Devices Garner Data on Insulation Performance; Smart Sensors Gather Information for Machine Diagnostics; Oxygen Sensors Monitor Bioreactors and Ensure Health and Safety; Vision Algorithms Catch Defects in Screen Displays; and Deformable Mirrors Capture Exoplanet Data, Reflect Lasers.

Source record↗

Investigating explainable transfer learning for battery lifetime prediction under state transitions

Battery lifetime prediction at early cycles is crucial for researchers and manufacturers to examine product quality and promote technology development. Machine learning has been widely utilized to construct data-driven solutions for high-accuracy predictions. However, the internal mechanisms of batteries are sensitive to many factors, such as charging/discharging protocols, manufacturing/storage conditions, and usage patterns. These factors will induce state transitions, thereby decreasing the prediction accuracy of data-driven approaches. Transfer learning is a promising technique that overcomes this difficulty and achieves accurate predictions by jointly utilizing information from various sources. Hence, we develop two transfer learning methods, Bayesian Model Fusion and Weighted Orthogonal Matching Pursuit, to strategically combine prior knowledge with limited information from the target dataset to achieve superior prediction performance. From our results, our transfer learning methods reduce root-mean-squared error by 41% through adapting to the target domain. Furthermore, the transfer learning strategies identify the variations of impactful features across different sets of batteries and therefore disentangle the battery degradation mechanisms and the root cause of state transitions from the perspective of data mining. These findings suggest that the transfer learning strategies proposed in our work are capable of acquiring knowledge across multiple data sources for solving specialized issues.

25 ENERGY STORAGE↗

Dimensionality Reduction Through Classifier Ensembles

In data mining, one often needs to analyze datasets with a very large number of attributes. Performing machine learning directly on such data sets is often impractical because of extensive run times, excessive complexity of the fitted model (often leading to overfitting), and the well-known "curse of dimensionality." In practice, to avoid such problems, feature selection and/or extraction are often used to reduce data dimensionality prior to the learning step. However, existing feature selection/extraction algorithms either evaluate features by their effectiveness across the entire data set or simply disregard class information altogether (e.g., principal component analysis). Furthermore, feature extraction algorithms such as principal components analysis create new features that are often meaningless to human users. In this article, we present input decimation, a method that provides "feature subsets" that are selected for their ability to discriminate among the classes. These features are subsequently used in ensembles of classifiers, yielding results superior to single classifiers, ensembles that use the full set of features, and ensembles based on principal component analysis on both real and synthetic datasets.

Oza, Nikunj C.↗

Precursor reaction pathway leading to BiFeO 3 formation: insights from text-mining and chemical reaction network analyses

BiFeO 3 (BFO) is a next-generation non-toxic multiferroic material with applications in sensors, memory devices, and spintronics, where its crystallinity and crystal structure directly influence its functional properties. Designing sol–gel syntheses that result in phase-pure BFO remains a challenge due to the complex interactions between metal complexes in the precursor solution. Here, we combine text-mined data and chemical reaction network (CRN) analysis to obtain novel insight into BFO sol–gel precursor chemistry. We perform text-mining analysis of 340 synthesis recipes with the emphasis on phase-pure BFO and identify trends in the use of precursor materials, including that nitrates are the preferred metal salts, 2-methoxyethanol (2 ME) is the dominant solvent, and adding citric acid as a chelating agent frequently leads to phase-pure BFO. Our CRN analysis reveals that the thermodynamically favored reaction mechanism between bismuth nitrate and 2ME interaction involves partial solvation followed by dimerization, contradicting assumptions in previous literature. We suggest that further oligomerization, facilitated by nitrite ion bridging, is critical for achieving the pure BFO phase.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

A Web-based Collaborative Tool for Mars Analog Data Exploration

Solving today's complex research and modeling challenges are dependent on our ability to discover, access, integrate, and share information from multiple sources. The planetary sciences community is no exception'; over the last few years, the need for data mining and exploration tools that can expedite comparative studies between Martian and terrestrial analogs sites and aid the interpretation of Mars data sets has become evident. Data sharing maximizes scientific return from studies and data sets.

Necsoiu, M.↗

NASA Stennis Space Center Integrated System Health Management Test Bed and Development Capabilities

Integrated System Health Management (ISHM) is a capability that focuses on determining the condition (health) of every element in a complex System (detect anomalies, diagnose causes, prognosis of future anomalies), and provide data, information, and knowledge (DIaK)-not just data-to control systems for safe and effective operation. This capability is currently done by large teams of people, primarily from ground, but needs to be embedded on-board systems to a higher degree to enable NASA's new Exploration Mission (long term travel and stay in space), while increasing safety and decreasing life cycle costs of spacecraft (vehicles; platforms; bases or outposts; and ground test, launch, and processing operations). The topics related to this capability include: 1) ISHM Related News Articles; 2) ISHM Vision For Exploration; 3) Layers Representing How ISHM is Currently Performed; 4) ISHM Testbeds & Prototypes at NASA SSC; 5) ISHM Functional Capability Level (FCL); 6) ISHM Functional Capability Level (FCL) and Technology Readiness Level (TRL); 7) Core Elements: Capabilities Needed; 8) Core Elements; 9) Open Systems Architecture for Condition-Based Maintenance (OSA-CBM); 10) Core Elements: Architecture, taxonomy, and ontology (ATO) for DIaK management; 11) Core Elements: ATO for DIaK Management; 12) ISHM Architecture Physical Implementation; 13) Core Elements: Standards; 14) Systematic Implementation; 15) Sketch of Work Phasing; 16) Interrelationship Between Traditional Avionics Systems, Time Critical ISHM and Advanced ISHM; 17) Testbeds and On-Board ISHM; 18) Testbed Requirements: RETS AND ISS; 19) Sustainable Development and Validation Process; 20) Development of on-board ISHM; 21) Taxonomy/Ontology of Object Oriented Implementation; 22) ISHM Capability on the E1 Test Stand Hydraulic System; 23) Define Relationships to Embed Intelligence; 24) Intelligent Elements Physical and Virtual; 25) ISHM Testbeds and Prototypes at SSC Current Implementations; 26) Trailer-Mounted RETS; 27) Modeling and Simulation; 28) Summary ISHM Testbed Environments; 29) Data Mining - ARC; 30) Transitioning ISHM to Support NASA Missions; 31) Feature Detection Routines; 32) Sample Features Detected in SSC Test Stand Data; and 33) Health Assessment Database (DIaK Repository).

Figueroa, Fernando↗

RNAV STAR Procedural Adherence

In this exploratory archival study we mined the performance of 24 major US airports area navigation standard terminal arrival routes (RNAV STARs) over the preceding three years. Overlaying radar track data on top of RNAV STAR routes provided a comparison between aircraft flight paths and the waypoint positions and altitude restrictions. NASA Ames Supercomputing resources were utilized to perform the data mining and processing. We investigated STARs by lateral transition path (full-lateral), vertical restrictions (full-lateral/full-vertical), and skipped waypoints (skips). In addition, we graphed altitudes and their frequencies of occurrence for altitude restrictions. Full-lateral compliance was generally greater than Full-lateral/full-vertical, but the delta between the rates was not always consistent. Full-lateral/full-vertical usage medians of the 2016 procedures ranged from 0 in KDEN (Denver) to 21 in KMEM (Memphis). Waypoint skips ranged from 0 to nearly 100 for specific waypoints. Altitudes restrictions were sometimes missed by systemic amounts in 1000 ft. increments from the restriction, creating multi-modal distributions. Other times, altitude misses looked to be more normally distributed around the restriction. This work is a preliminary investigation into the objective performance of instrument procedures and provides a framework to track how procedural concepts and design intervention function. In addition, this tool may aid in providing acceptability metrics as well as risk assessment information.

Stewart, Michael J.↗

Examining Runner’s Outdoor Heat Exposure Using Urban Microclimate Modeling and GPS Trajectory Mining

It is important to quantify human heat exposure in order to evaluate and mitigate the negative impacts of heat on human well-being in the context of global warming. This study proposed a human-centric framework to examine human personal heat exposure based on anonymous GPS trajectories data mining and urban microclimate modeling. The mean radiant temperature (Tmrt) that represents the human body’s energy balance was used to indicate human heat exposure. The meteorological data and high-resolution 3D urban model generated from multispectral remotely sensed images and LiDAR data were used as inputs in urban microclimate modeling to map the spatio-temporal distribution of the Tmrt in the Boston metropolitan area. The anonymous human GPS trajectory data collected from fitness Apps was used to map the spatiotemporal distribution of human outdoor activities. By overlaying the anonymous GPS trajectories on the generated spatio-temporal maps of Tmrt, this study further examined the heat exposure of runners in different age-gender groups in the Boston area. Results show that there is no significant difference in terms of heat exposure for female and male runners. The female runners in the age of 45-54 are exposed to more heat than female runners of 18-24 and 25-34, while there is no significant difference among male runners. This study proposed a novel method to estimate human heat exposure, which would shed new light on mitigating the negative impacts of heat on human health.

Personal heat exposure↗

A neural network for determination of latent dimensionality in Nonnegative Matrix Factorization

Non-negative Matrix Factorization (NMF) has proven to be a powerful unsupervised learning method for uncovering hidden features in complex and noisy datasets with applications in data mining, text recognition, dimension reduction, face recognition, anomaly detection, blind source separation, and many other fields. An important input for NMF is the latent dimensionality of the data, that is, the number of hidden features, K, present in the explored dataset. Unfortunately, and this quantity is rarely known a priori. The existing methods for determining latent dimensionality, such as Automatic Relevance Determination (ARD), are mostly heuristic and utilize different characteristics to estimate the number of hidden features. However, all of them require human presence to make a final determination of K. Here we utilize a supervised machine learning approach in combination with a recent method for model determination, called NMFk, to determine the number of hidden features automatically. NMFk performs a set of NMF simulations on an ensemble of matrices, obtained by bootstrapping the initial dataset, and estimates which K produces stable groups of latent features that reconstruct the initial dataset well. We then train a Multi-Layter Perceptron (MLP) classifier network to determine the correct number of latent features utilizing the statistics and characteristics of the NMF solution, obtained from NMFk. In order to train the MLP classifier, a training set of 58,660 matrices with predetermined latent features were factorized with NMFk. The MLP classifier in conjunction with NMFk maintains a greater than 95% success rate when applied to a held out test set. Additionally, when applied to two well-known benchmark datasets, the swimmer and MIT face data, NMFk/MLP correctly recovers the established number of hidden features. Finally, we compare the accuracy of our method to the ARD, AIC and Stability-based methods.

97 MATHEMATICS AND COMPUTING↗