Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “feature learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

SAXS Assistant: Automated SAXS analysis for structural discovery in biologics and polymeric nanoparticles

Small-angle x-ray scattering (SAXS) is a powerful technique for assessing macromolecular structure. High-throughput SAXS is limited by the time-consuming and, at times, subjective nature of SAXS data interpretation. Here, we present SAXS Assistant, a Python-based script that streamlines SAXS data analysis to extract features for machine learning (ML) and key structural parameters, including the Guinier radius of gyration (R g ), pair distance distribution function (PDDF)-derived R g , maximum particle dimension (D max ), and Kratky plots. The script builds upon BioXTAS RAW and validates reliability via Guinier/PDDF R g agreement, an important indicator of well-measured data sets. For assistance in D max estimation, a multilayer perceptron regressor was trained with 1940 data files from the Small Angle Scattering Biological Data Bank. The model achieved a test set performance R 2 = 0.90 and mean absolute error = 11.7 Å. Training exclusively with experimental data translates analyses from researchers, including experts in the field, to the ML model, which helps assess D max estimations from PDDF. Gaussian mixture model clustering was implemented to classify profiles into structural classes based on entries in the Small Angle Scattering Biological Data Bank. Users may therefore assess the similarity between experimental samples and known biomolecular shapes within the mapped repository entries. This probabilistic clustering aids in quantifying information from Kratky and generating shape-descriptive features. SAXS Assistant accelerates SAXS data analysis through enforced quality control, ML-ready outputs, and flags for low-confidence results. In addition to providing the ability to analyze large data sets at high throughput, this tool is versatile and may serve researchers in both biological and synthetic polymer research fields.

36 MATERIALS SCIENCE↗

Understanding the Impact of Unobservable Variables on the Performance of Predictive Models: The Need for Feature Space Partitioning and Fusion

When developing predictive models over a dataset, the model is globally optimized across the entire feature space to learn a decision boundary. However, when unobservable variables—which cannot be measured or estimated—interact with the observable variables, this can negatively impact the optimization applied to the decision boundary since the data samples introduced by unobservable variables may have little to no association with the applied global optimization. This, consequently, penalizes the entire decision boundary and model performance. This paper examines some of the detrimental effects of unobservable variables, particularly their role in creating new modes in the distribution of observable variables and reducing the separability of class distributions. Such challenges result in skewed or warped decision boundaries and decreased accuracy of model predictions, particularly for interpretable models like logistic regression and decision trees. Through two illustrative case examples, we highlight the need to address the challenges imposed by unobservable variables. We propose a strategy to mitigate these challenges by creating local regions within the feature space through partitioning. This enables the optimization of local models within the regions to overcome the impact of unobservability in different feature space localities. Research into a more sophisticated partitioning strategy and where the partition should be relative to the sample of interest is left as future work. Through the analysis of the impact of unobservability and the development of a partitioning method, we demonstrate the clear need for a partitioning strategy that integrates knowledge from multiple local models to estimate risk factors using information fusion. Thus, we establish the foundation and motivation for using partitioning and information fusion to overcome the effects of unobservability in predictive models. Formal fusion methods, such as Dempster-Shafer theory, can better leverage the information from local regions to improve the performance of interpretable predictive models in the presence of unobservable variables.

Time Series Data↗

U-Surf: a global 1 km spatially continuous urban surface property dataset for kilometer-scale urban-resolving Earth system modeling

High-resolution urban climate modeling has faced substantial challenges due to the absence of a globally consistent, spatially continuous, and accurate dataset to represent the spatial heterogeneity of urban surfaces and their biophysical properties. This deficiency has long obstructed the development of urban-resolving Earth system models (ESMs) and ultra-high-resolution urban climate modeling, over large domains. Here, we present U-Surf, a first-of-its-kind 1 km resolution present-day (circa 2020) global continuous urban surface parameter dataset. Using the urban canopy model (UCM) in the Community Earth System Model as a base model for satisfying dataset requirements, U-Surf leverages the latest advances in remote sensing, machine learning, and cloud computing to provide the most relevant urban surface biophysical parameters, including radiative, morphological, and thermal properties, for UCMs at the facet and canopy level. Generated using a systematically unified workflow, U-Surf ensures internal consistency among key parameters, making it the first globally coherent urban canopy surface dataset. U-Surf significantly improves the representation of the urban land heterogeneity both within and across cities globally; provides essential, high-fidelity surface biophysical constraints to urban-resolving ESMs; enables detailed city-to-city comparisons across the globe; and supports next-generation kilometer-resolution Earth system modeling across scales. U-Surf parameters can be easily converted or adapted to various types of UCMs, such as those embedded in weather and regional climate models, as well as air quality models. The fundamental urban surface constraints provided by U-Surf can also be used as features for machine learning models and can have other broad-scale applications for socioeconomic, public health, and urban planning contexts. We expect U-Surf to advance the research frontier of urban system science, climate-sensitive urban design, and coupled human–Earth systems in the future. The dataset is publicly available at https://doi.org/10.5281/zenodo.11247598 (Cheng et al., 2024).

Cheng, Yifan [Univ. of Illinois at Urbana-Champaig↗

Aging matrix visualizes complexity of battery aging across hundreds of cycling protocols

To reliably deploy lithium-ion batteries, a fundamental understanding of cycling aging behavior is critical. Battery aging consists of complex and highly coupled phenomena, making it challenging to develop a holistic interpretation. In this work, we generate a diverse battery cycling dataset with a broad range of degradation trajectories, consisting of 359 high energy density commercial Li(Ni,Co,Al)O 2 /graphite + SiO x cylindrical 21 700 cells cycled across 207 unique cycling protocols. We consolidate aging via 16 mechanistic state-of-health (SOH) metrics, including cell-level performance metrics, electrode-specific capacities/state-of-charges (SOCs), and aging trajectory metrics. We develop a framework using interpretable machine learning and explainable features to generate an aging matrix that visually deconvolutes the complex battery degradation behavior. This generalizable data-driven mechanistic framework simplifies the complex interplay between cycling conditions, degradation modes, and SOH, acting as a hypothesis-generation tool to aid battery users in identifying key degradation regimes for further study and experimentation.

25 ENERGY STORAGE↗

Breaking the barrier of human-annotated training data for machine learning-aided plant research using aerial imagery

Machine learning (ML) can accelerate biological research. However, the adoption of such tools to facilitate phenotyping based on sensor data has been limited by (i) the need for a large amount of human-annotated training data for each context in which the tool is used and (ii) phenotypes varying across contexts defined in terms of genetics and environment. This is a major bottleneck because acquiring training data is generally costly and time-consuming. This study demonstrates how a ML approach can address these challenges by minimizing the amount of human supervision needed for tool building. A case study was performed to compare ML approaches that examine images collected by an uncrewed aerial vehicle to determine the presence/absence of panicles (i.e. “heading”) across thousands of field plots containing genetically diverse breeding populations of 2 Miscanthus species. Automated analysis of aerial imagery enabled the identification of heading approximately 9 times faster than in-field visual inspection by humans. Leveraging an Efficiently Supervised Generative Adversarial Network (ESGAN) learning strategy reduced the requirement for human-annotated data by 1 to 2 orders of magnitude compared to traditional, fully supervised learning approaches. The ESGAN model learned the salient features of the data set by using thousands of unlabeled images to inform the discriminative ability of a classifier so that it required minimal human-labeled training data. This method can accelerate the phenotyping of heading date as a measure of flowering time in Miscanthus across diverse contexts (e.g. in multistate trials) and opens avenues to promote the broad adoption of ML tools.

59 BASIC BIOLOGICAL SCIENCES↗

Evidence for B + → K + ν ν ¯ decays

We search for the rare decay B + → K + ν ν ¯ in a 362 fb − 1 sample of electron-positron collisions at the ϒ ( 4 S ) resonance collected with the Belle II detector at the SuperKEKB collider. We use the inclusive properties of the accompanying B meson in ϒ ( 4 S ) → B B ¯ events to suppress background from other decays of the signal B candidate and light-quark pair production. We validate the measurement with an auxiliary analysis based on a conventional hadronic reconstruction of the accompanying B meson. For background suppression, we exploit distinct signal features using machine learning methods tuned with simulated data. The signal-reconstruction efficiency and background suppression are validated through various control channels. The branching fraction is extracted in a maximum likelihood fit. Our inclusive and hadronic analyses yield consistent results for the B + → K + ν ν ¯ branching fraction of [ 2.7 ± 0.5 ( stat ) ± 0.5 ( syst ) ] × 10 − 5 and [ 1.1 − 0.8 + 0.9 ( stat ) − 0.5 + 0.8 ( syst ) ] × 10 − 5 , respectively. Combining the results, we determine the branching fraction of the decay B + → K + ν ν ¯ to be [ 2.3 ± 0.5 ( stat ) − 0.4 + 0.5 ( syst ) ] × 10 − 5 , providing the first evidence for this decay at 3.5 standard deviations. The combined result is 2.7 standard deviations above the standard model expectation. Published by the American Physical Society 2024

Astronomy & Astrophysics↗

pixelvar79/ESGAN-Flowering-Detection-paper

Machine learning (ML) can accelerate biological research. However, the adoption of such tools to facilitate phenotyping based on sensor data has been limited by (i) the need for a large amount of human-annotated training data for each context in which the tool is used and (ii) phenotypes varying across contexts defined in terms of genetics and environment. This is a major bottleneck because acquiring training data is generally costly and time-consuming. This study demonstrates how a ML approach can address these challenges by minimizing the amount of human supervision needed for tool building. A case study was performed to compare ML approaches that examine images collected by an uncrewed aerial vehicle to determine the presence/absence of panicles (i.e. “heading”) across thousands of field plots containing genetically diverse breeding populations of 2 Miscanthus species. Automated analysis of aerial imagery enabled the identification of heading approximately 9 times faster than in-field visual inspection by humans. Leveraging an Efficiently Supervised Generative Adversarial Network (ESGAN) learning strategy reduced the requirement for human-annotated data by 1 to 2 orders of magnitude compared to traditional, fully supervised learning approaches. The ESGAN model learned the salient features of the data set by using thousands of unlabeled images to inform the discriminative ability of a classifier so that it required minimal human-labeled training data. This method can accelerate the phenotyping of heading date as a measure of flowering time in Miscanthus across diverse contexts (e.g. in multistate trials) and opens avenues to promote the broad adoption of ML tools.

Varela, Sebastian↗

Co-Simulation Meets AI: MCP-Driven Power System Analysis

GridGPT, a fine-tuned Generative AI model is designed for on-premise use in grid control rooms. This presentation will demonstrate how eGridGPT can seamlessly integrate with control room solutions to offer operators, engineers, and corporate users enhanced guidance and decision support. It is to show how this innovative AI solution can improve state estimation, boost variable energy forecasting, and optimize grid operations. By leveraging eGridGPT's unique features, audience will learn to unlock new levels of automation, predictive analytics, and reliability within their power systems, ultimately leading to reduced downtime and improved operational efficiency.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Unmanned Aircraft Systems (UAS) and Light Detection and Ranging (LiDAR)/Camera Technologies to Detect Avian Events and Other Environmental Measures at Utility- Scale Power Plants (Final Report)

The goal of this project was to develop and validate two complementary, cost-effective remote sensing technologies to monitor avian fatalities at utility-scale solar facilities: fixed platform (Animal Activity Monitoring-AAM) and aerial-based (Uncrewed Aircraft Systems-UAS). This project used these features with machine learning to automate the detection of avian carcasses and nests at solar facilities.

14 SOLAR ENERGY↗

Effects of lateral resolution on the identification of volcanotectonic provinces on earth and Venus

In an attempt to learn what volcanotectonic features can still be discerned in continental and oceanic areas of the earth when topographic data are degraded to simulate the data sampled by the Pioneer-Venus altimeter, two digital topographic data sets (the 30-second continental U.S. altitude data and the 30 x 30 nautical mile bathymetry data for the North Pacific) were degraded and displayed in the same way as the altimeter data from Venus. The Appalachians were reduced to a gentle swell, with a wavelength of 300 km, and a height of 500 m. The Cordillera was seen as a broad swell, 2500 km wide, and about 2 km high. The east Pacific rise, east Pacific fractures, seamount chains, the Hawaiian swell, and most trenches were discernible in the degraded Pacific data; whereas rises, transforms, seamount chains, and trenches were not seen in the Venus data, even after corrections were made for the higher surface temperature and the absence of oceans on Venus. It was concluded that a plate tectonic regime, similar to earth's does not currently appear to exist on Venus. As shown by the Cordillera data, the Pioneer-Venus information is not of sufficiently high quality to discern whether the highlands of Venus preserve evidence for orogenic events related to plate tectonics.

Arvidson, R. E.↗

A composite self tuning strategy for fuzzy control of dynamic systems

The feature of self learning makes fuzzy logic controllers attractive in control applications. This paper proposes a strategy to tune the fuzzy logic controller on-line by tuning the data base as well as the rule base. The structure of the controller is outlined and preliminary results are presented using simulation studies.

Shieh, C.-Y.↗

Content Documents Management

The Content Documents are created and managed under the System Software group with. Launch Control System (LCS) project. The System Software product group is lead by NASA Engineering Control and Data Systems branch (NE~C3) at Kennedy Space Center. The team is working on creating Operating System Images (OSI) for different platforms (i.e. AIX, Linux, Solaris and Windows). Before the OSI can be created, the team must create a Content Document which provides the information of a workstation or server, with the list of all the software that is to be installed on it and also the set where the hardware belongs. This can be for example in the LDS, the ADS or the FR-l. The objective of this project is to create a User Interface Web application that can manage the information of the Content Documents, with all the correct validations and filters for administrator purposes. For this project we used one of the most excellent tools in agile development applications called Ruby on Rails. This tool helps pragmatic programmers develop Web applications with Rails framework and Ruby programming language. It is very amazing to see how a student can learn about OOP features with the Ruby language, manage the user interface with HTML and CSS, create associations and queries with gems, manage databases and run a server with MYSQL, run shell commands with command prompt and create Web frameworks with Rails. All of this in a real world project and in just fifteen weeks!

Muniz, R.↗

DRAGON - 8U Nanosatellite Orbital Deployer

The Space Research Centre of the Polish Academy of Sciences (SRC PAS) together with Astronika company have developed an Orbital Deployer called DRAGON for ejection of the Polish scientific nanosatellite BRITE-PL Heweliusz (Fig. 1). The device has three unique mechanisms including an adopted and scaled lock and release mechanism from the ESA Rosetta mission MUPUS instrument. This paper discusses major design restrictions of the deployer, unique design features, and lessons learned from development through testing.

Dobrowolski, Marcin↗

The Right Amount of Glue: Technologies and Standards Relevant to a Future Solar-Terrestrial Data Environment

In order to meet the challenge of developing a new system science, we will need to employ technology that enables researchers to access data from fields with which they are at least initially unfamiliar as well as from sources they use more regularly. At the same time, the quantity of data to be obtained by missions such as the Solar Dynamics Observatory demands ease and simplicity of data access. These competing demands must in turn fit within severely constrained funding for data analysis in such projects. Based on experience in only a single discipline but with a diversity of data types and sources, we will give examples of technology that have made a significant difference in the way people do science. Similarly, we will show how adoption of a well-documented data format has made it easier for one community to search, reduce, and analyze data. We will also describe a community-supported data reduction and analysis software tree with useful features. We will attempt to generalize the lessons learned in these instances to features the broader, solar-terrestrial community might find compelling, while avoiding overdesign of a common data environment.

Gurman, J. B.↗

Interpretable Machine Learning for Molecular Biosignatures: a Novel Single-Sample Feature Importance Method That Is Sensitive To Statistical Interactions

Isotope ratio mass spectrometry (IRMS) of volatiles (e.g., CO 2 ) promises to be a powerful tool for potential biosignature detection for future missions to ocean worlds (OW) such as Europa and Enceladus. Machine learning (ML) methods for IRMS data could enable science autonomy by onboard prediction of seawater chemistry and biosignature presence. However, ML models are likely to be complex and involve statistical interactions between features (variables), which can make predictions seem opaque and enigmatic. For ML predictions as significant as extraterrestrial biosignatures, we must place extraordinary confidence in models. It is therefore essential that these models make interpretable predictions (i.e., human-understandable) and include false-prediction diagnostics. We achieve high accuracy and interpretability in ML biosignature and seawater chemistry models for OW through a nearest-neighbors feature selection tool that detects statistical interactions between predictors, constructs interaction networks for visualization of selected features working together to make a prediction, and reports single-sample feature importance scores for false-detection diagnostics. Here we develop a novel single-sample nearest-neighbors projected distance regression(ssNPDR) feature selection method that improves upon existing single-sample algorithms through the inclusion of statistical interactions while providing false-prediction diagnostics for ML models.

geochemistry↗

Physics-informed latent neural operator for real-time predictions of time-dependent parametric PDEs

Deep operator network (DeepONet) has shown significant promise as surrogate models for systems governed by partial differential equations (PDEs), enabling accurate mappings between infinite-dimensional function spaces. However, when applied to systems with high-dimensional input-output mappings arising from large numbers of spatial and temporal collocation points, these models often require heavily overparameterized networks, leading to long training times. Latent DeepONet addresses some of these challenges by introducing a two-step approach: first learning a reduced latent space using a separate model, followed by operator learning within this latent space. While efficient, this method is inherently data-driven and lacks mechanisms for incorporating physical laws, limiting its robustness and generalizability in data-scarce settings. Here, in this work, we propose PI-Latent-NO, a physics-informed latent neural operator framework that integrates governing physics directly into the learning process. Our architecture features two coupled DeepONets trained end-to-end: a Latent-DeepONet that learns a low-dimensional representation of the solution, and a Reconstruction-DeepONet that maps this latent representation back to the physical space. By embedding PDE constraints into the training via automatic differentiation, our method eliminates the need for labeled training data and ensures physics-consistent predictions. The proposed framework is both memory and compute-efficient, exhibiting near-constant scaling with problem size and demonstrating significant speedups over traditional physics-informed operator models. We validate our approach on a range of parametric PDEs, showcasing its accuracy, scalability, and suitability for real-time prediction in complex physical systems.

Latent representations↗

Understanding Mixed Precision GEMM with MPGemmFI: Insights into Fault Resilience

Emerging deep learning workloads urgently need fast general matrix multiplication (GEMM). Thus, one of the critical features of machine-learning-specific accelerators such as NVIDIA Tensor Cores, AMD Matrix Cores, and Google TPUs is the support of mixed-precision enabled GEMM. For DNN models, lower-precision FP data formats and computation offer acceptable correctness but significant performance, area, and memory footprint improvement. While promising, the mixed-precision computation on error resilience remains unexplored. To this end, we develop a fault injection framework that systematically injects fault into the mixed-precision computation results. We investigate how the faults affect the accuracy of machine learning applications. Based on the characteristics of error resilience, we offer lightweight error detection and correction solutions that significantly improve the overall model accuracy by 75% if the models experience hardware faults. The solutions can be efficiently integrated into the accelerator's pipelines.

Fang, Bo↗

Remote Sensing in Archaeology: Visible Temporal Change of Archaeological Features of the Peten, Guatemala

The purpose of this archaeological research was two-fold; the location of Mayan sites and features in order to learn more of this cultural group, and the (cultural) preservation of these sites and features for the future using Landsat Thematic Mapper (TM) images. Because the rainy season, traditionally at least, lasts about six months (about June to December), the time of year the image is acquired plays an important role in spectral reflectance. Images from 1986, 1995, and 1997 were selected because it was felt they would provide the best opportunity for success in layering different bands from different years together to attempt to see features not completely visible in any one year. False-color composites were created including bands 3, 4, and 5 using a mixture of years and bands. One particular combination that yielded tremendously interesting results included band 5 from 1997, band 4 from 1995, and band 3 from 1986. A number of straight linear features (probably Mayan causeways) run through the bajos that Dr. Sever believes are features previously undiscovered. At this point, early indications are that this will be a successful method for locating "new" Mayan archaeological features in the Peten.

Lowry, James D., Jr.↗