Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Extreme learning machine”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Multi-Class Anomaly Detection in Flight Data using Semi-Supervised Explainable Deep Learning Model

Identifying precursor for safety incidents in aviation data is a crucial task, yet extremely challenging. The main approach, in practice, leverages domain expertise to define expected tolerances in system’s behavior and alarm exceedance from such safety margins. However, this approach is incapable of identifying unknown risk and vulnerabilities. Machine learning has been long studied and deployed to identify precursors for such anomalies, with the great challenge of the need for sufficient labelled set of data to achieve a reliable and accurate performance. In this article, we develop an explainable deep semi-supervised model for anomaly detection in aviation, building upon recent advancements in the machine learning literature. The proposed model combines feature engineering and classification in the feature space, while leveraging all available data (labelled and unlabeled). Validating on two case studies of anomaly detection in take-off and landing phases of commercial aircraft, we show that our model is able to outperform state-of-the-art supervised anomaly detection model and reach significantly high accuracy and low false alarm with minimum amount of available labelled data.

Anomaly Detection↗

A New Look at NASA: Strategic Research In Information Technology

This viewgraph presentation provides information on research undertaken by NASA to facilitate the development of information technologies. Specific ideas covered here include: 1) Bio/nano technologies: biomolecular and nanoscale systems and tools for assembly and computing; 2) Evolvable hardware: autonomous self-improving, self-repairing hardware and software for survivable space systems in extreme environments; 3) High Confidence Software Technologies: formal methods, high-assurance software design, and program synthesis; 4) Intelligent Controls and Diagnostics: Next generation machine learning, adaptive control, and health management technologies; 5) Revolutionary computing: New computational models to increase capability and robustness to enable future NASA space missions.

Alfano, David↗

Landslide Likelihood Prediction using Machine Learning Algorithms

The supply of electricity via power plants is criticalto the operation of many critical infrastructure systems in mod-ern society. Natural hazards can disrupt the power supply, causepower outages that can halt economic growth, and impede emer-gency response until power is restored. The proposed work aimsto predict the landslides likelihood in these critical infrastructurelocations in the Northeastern USA using integrated databases ofexplanatory variables and machine learning algorithms. First,data related to landslides are obtained and merged, includingtopographic, soil moisture, and precipitation-related data. Fiveregression algorithms, namely: Random Forest, Extreme Gradi-ent Boosting (XGBoost), K-Nearest Neighbor regression (KNN),Linear Support Vector Regressor (SVR), and Linear regression,are utilized to predict the landslide probability and evaluatedon the dataset. The accuracy of the models is assessed by usingstatistical metrics such as mean absolute error (MAE), meansquared error (MSE), and root mean squared error (RMSE).The study results show that Random Forest outperformed othermodels with the mutual information feature selection method.It achieved an MSE of 0.0011 with mutual information-basedfeature selection and an MSE of 0.00157 without feature selection.KNN regressor outperformed the other models with an MSEof 0.00139 with correlation-based information selection. Theproposed landslide identification model with Random Forestalgorithm shows outstanding robustness and great potential intackling the landslide likelihood prediction by employing MLalgorithms.

Vasundhara Acharya↗

Science Autonomy for Ocean Worlds Astrobiology: A Perspective

Astrobiology missions to ocean worlds in our solar system must overcome both scientific and technological challenges due to extreme temperature and radiation conditions, long communication times, and limited bandwidth. While such tools could not replace ground-based analysis by science and engineering teams, machine learning algorithms could enhance the science return of these missions through development of autonomous science capabilities. Examples of science autonomy include onboard data analysis and subsequent instrument optimization, data prioritization (for transmission), and real-time decision-making based on data analysis. Similar advances could be made to develop streamlined data processing software for rapid ground-based analyses. Here we discuss several ways machine learning and autonomy could be used for astrobiology missions, including landing site selection, prioritization and targeting of samples, classification of “features” (e.g., proposed biosignatures) and novelties (uncharacterized, “new” features, which may be of most interest to agnostic astrobiological investigations), and data transmission.

ocean worlds↗

A Strategic Approach for Dense, Integrated, Vehicle Navigation

Drone usage has been on the rise in recent years with applications that include parcel delivery, wildlife protection, precision farming, law enforcement, and industrial inspection, just to name a few. Once regulations and safety policies are put in place to allow for the widespread use of unmanned drones, the number of aircraft in the National Airspace System (NAS) is expected to skyrocket to millions, potentially congesting the airspace which increases the likelihood of separation violations and possibly incidents. Currently, flight infrastructure can only support a few thousand aircraft flying over the United States National Airspace System (NAS) at any given time. A delay at one airport can send ripple effects throughout the system, causing more delays and missed connections. In air traffic control, separation is the concept of keeping an “ownship” aircraft outside a minimum distance from “intruder” aircraft to reduce the risk of the aircraft colliding, as well as preventing accidents due to secondary factors, such as wake turbulence. Maintaining proper separation is often a safety critical property for fixed-wing drones in the airspace. This paper addresses drone separation in time and distance for high volume corridors (en route) and lanes (on ground), merging as well as crossing intersections of multiple corridors/lanes. In this paper, the term drone is applied to both Unmanned Aerial Vehicle (UAV) and small Unmanned Aircraft System (UAS) vehicles operating autonomously. There exists a gamut of approaches to the merging and crossing problems. At one end of the extreme are the conservative yet low cost and verifiable solutions of today that deal with two drones at a time. At the other end are complex Machine Learning-based solutions with high computing requirements for fully autonomous drones of the future that are expected to handle all contentions. This paper presents a feasible and verifiable strategic approach to these problems that is based on distributed cooperation between the drones and the infrastructure. Three phases of the strategic approach (Prepare, Adjust, Commit) are presented. Simulation results are presented that show the proposed approach is stable and resilient to induced perturbations and guarantees a set of fixed-wing drones to merge and cross intersections by adjusting their speed based on their distance to the aircraft in front of them while remaining in the equilibrium state. The equilibrium state is defined as the state when a set of n aircraft move at a relatively constant speed and uniform spacing from each other in a congested system. A congested system is defined as the state when at least one aircraft cannot move at its maximum allowed speed. Unlike existing centralized and pre-planned approaches, the proposed solution is fully distributed and enables autonomous aircraft to decide to adjust their speed and distance with respect to the preceding aircraft, dynamically. Simulation results are presented that assess the feasibility of the approach.

Distributed↗

A Strategic Approach for Dense, Integrated, Vehicle Navigation

Drone usage has been on the rise in recent years with applications that include parcel delivery, wildlife protection, precision farming, law enforcement, and industrial inspection, just to name a few. Once regulations and safety policies are put in place to allow for the widespread use of unmanned drones, the number of aircraft in the National Airspace System (NAS) is expected to skyrocket to millions, potentially congesting the airspace which increases the likelihood of separation violations and possibly incidents. Currently, flight infrastructure can only support a few thousand aircraft flying over the United States National Airspace System (NAS) at any given time. A delay at one airport can send ripple effects throughout the system, causing more delays and missed connections. In air traffic control, separation is the concept of keeping an “ownship” aircraft outside a minimum distance from “intruder” aircraft to reduce the risk of the aircraft colliding, as well as preventing accidents due to secondary factors, such as wake turbulence. Maintaining proper separation is often a safety critical property for fixed-wing drones in the airspace. This paper addresses drone separation in time and distance for high volume corridors (en route) and lanes (on ground), merging as well as crossing intersections of multiple corridors/lanes. In this paper, the term drone is applied to both Unmanned Aerial Vehicle (UAV) and small Unmanned Aircraft System (UAS) vehicles operating autonomously. There exists a gamut of approaches to the merging and crossing problems. At one end of the extreme are the conservative yet low cost and verifiable solutions of today that deal with two drones at a time. At the other end are complex Machine Learning-based solutions with high computing requirements for fully autonomous drones of the future that are expected to handle all contentions. This paper presents a feasible and verifiable strategic approach to these problems that is based on distributed cooperation between the drones and the infrastructure. Three phases of the strategic approach (Prepare, Adjust, Commit) are presented. Simulation results are presented that show the proposed approach is stable and resilient to induced perturbations and guarantees a set of fixed-wing drones to merge and cross intersections by adjusting their speed based on their distance to the aircraft in front of them while remaining in the equilibrium state. The equilibrium state is defined as the state when a set of n aircraft move at a relatively constant speed and uniform spacing from each other in a congested system. A congested system is defined as the state when at least one aircraft cannot move at its maximum allowed speed. Unlike existing centralized and pre-planned approaches, the proposed solution is fully distributed and enables autonomous aircraft to decide to adjust their speed and distance with respect to the preceding aircraft, dynamically. Simulation results are presented that assess the feasibility of the approach.

Distributed↗

A Strategic Approach for Dense, Integrated, Vehicle Navigation

Drone usage has been on the rise in recent years with applications that include parcel delivery, wildlife protection, precision farming, law enforcement, and industrial inspection, just to name a few. This paper addresses drone separation in time and distance for high volume corridors (en route) and lanes (on ground), merging as well as crossing intersections of multiple corridors/lanes. There exists a gamut of approaches to solving merging and intersection crossing problems. At one end of the extreme are the conservative yet low cost and verifiable solutions of today that deal with two drones at a time. At the other end are complex Machine Learning-based solutions with high computing requirements for fully autonomous drones of the future that are expected to handle all contentions. This paper presents a feasible and verifiable strategic approach to solving these problems that is based on distributed cooperation between the UAVs/UASs and the infrastructure. Unlike existing centralized and pre-planned approaches, the proposed solution is fully distributed and enables autonomous aircraft to decide to adjust their speed and distance with respect to the preceding aircraft, dynamically. Three phases of the strategic approach (Prepare, Adjust, Commit) are presented. Simulation results are presented that show the proposed approach is stable and resilient to induced perturbations and guarantees a set of fixed-wing UAVs/UASs to merge and cross intersections by adjusting their speed based on their distance to the aircraft in front of them.

Distributed↗

Pushing the Limits of Aquatic Remote Sensing: Synthetic Data and Deep Learning for Fast Inverse Emulation of A Coupled Ocean-Atmosphere Radiative Transfer Model

The inversion of electromagnetic information to physical and biological properties of the water column is a notoriously difficult problem, yet fundamental to our ability of understanding aquatic processes on large time and space scales. There is now a growing necessity to develop pragmatic approaches that allow timely and effective extrapolation of local processes, to spatially resolved global products, and to promote operational and sustainable resource policy management. This presentation will discuss research integrating advanced biological and radiative modeling, high-end computation, and machine learning to develop a portable global processor for simultaneous retrieval of atmosphere and water optics for diverse aquatic systems from the open and coastal ocean to optically extreme inland waters and harmful algal blooms. We will discuss some of the basic concepts behind the forward modeling approach including DEAP, the novel Distributed Equivalent Algal Populations model, for developing large spectral libraries of aquatic particle optics to aid in our ability to distinguish phytoplankton functional types (PFTs) and inorganic material, as well as other factors which enable comprehensive modeling from the benthos to top-of-atmosphere (TOA). This information is being used to understand how we can leverage next-generation deep learning methods for maximum information retrieval and rapid image processing, while also providing capabilities to identify minimum sensor spectral requirements necessary for certain aquatic applications. Further, I will touch on how we envision this research to enable the aquatic community for science discovery and how we are moving closer towards the capability for high-fidelity global analysis of aquatic ecosystems.

Jeremy Alan Kravitz↗

Machine Learning Emulators and Empirical Models Combining Climate and Global Crop Models for Seasonal Agricultural Production

We present results from several connected efforts to apply machine learning methods to estimates of seasonal agricultural production anomalies around the world. First, we apply the XGBoost Random Forest method to fit emulators that mimic global crop models participating in the Agricultural Model Intercomparison and Improvement Project (AgMIP) Global Gridded Crop Model Intercomparison (GGCMI). These are the same models used in the agricultural sector simulations of the Inter-Sectoral Impacts Model Intercomparison Project (ISIMIP). These emulators use 8 climate variables split across 5 sub-seasonal representations of the growing season for each ½ degree grid cell around the world for maize, wheat, rice and soybeans. Emulators are useful for estimating conditions that have not already been simulated by GGCMI (e.g., in a seasonal prediction model) and also to diagnose model differences and capabilities. For example, emulators of the pDSSAT maize model tend to be more reliant on mean temperatures than the LPJmL model, and few models have strong responses to cold extremes. Second, we use a similar XGBoost approach to fit empirical models for national production data for the top 20 producing countries according to the United Nations Food and Agricultural Organization (FAO). Models utilize both climate observations and the GGCM models as predictors, resulting in skillful models for many (but not all) top producing-countries. The patterns of climate and crop model features selected indicate regions and systems that are better or worse simulated by the GGCMs. For example, information in cold extreme predictors is often combined with GGCM output predictors to provide sensitivity that models may underrepresent.

machine learning↗

Reconstructing PM 2.5 Data Record for the Kathmandu Valley Using a Machine Learning Model

This paper presents a method for reconstructing the historical hourly concentrations of particulate matter 2.5 (PM2.5) over the Kathmandu Valley from 1980 to the present. The method uses a machine learning model that is trained using PM2.5 readings from US Embassy (Phora Durbar) as a ground truth, and the meteorological data from Modern-Era Retrospective Analysis for Research and Applications v2 (MERRA2) as input. The Extreme Gradient Boosting (XGBoost) model acquires a credible 10-fold cross-validation (CV) score of ~83.4%, an r2-score of ~84%, a Root Mean Square Error (RMSE) of ~15.82 µg/m3, and a Mean Absolute Error (MAE) of ~10.27 µg/m3. Further demonstrating the model's applicability to years other than those for which truth values are unavailable, the multiple cross-test with an unseen data set offered r2-scores for 2018, 2019, and 2020 ranging from 56% to 67%. The model-predicted data agrees with true values and indicates that MERRA2 underestimates PM2.5 over the region. It strongly agrees with ground-based evidence showing substantially higher mass concentrations in the dry pre- and post-monsoon seasons than in the monsoon months. It also shows a strong anti-correlation between PM2.5 concentration and humidity. The results also demonstrate that none of the years fulfilled the annual mean air quality index (AQI) standards set by the World Health Organization (WHO).

machine learning↗

Vegetation Greening Mitigates the Impacts of Increasing Extreme Rainfall on Runoff Events

Future flood risk assessment has primarily focused on heavy rainfall as the main driver, with the assumption that projected increases in extreme rain events will lead to subsequent flooding. However, the presence of and changes in vegetation have long been known to influence the relationship between rainfall and runoff. Here, we extract historical (1850–1880) and projected (2070–2100) daily extreme rainfall events, the corresponding runoff, and antecedent conditions simulated in a prominent large Earth system model ensemble to examine the shifting extreme rainfall and runoff relationship. Even with widespread projected increases in the magnitude (78% of the land surface) and number (72%) of extreme rainfall events, we find projected declines in event-based runoff ratio (runoff/rainfall) for a majority (57%) of the Earth surface. Runoff ratio declines are linked with decreases in antecedent soil water driven by greater transpiration and canopy evaporation (both linked to vegetation greening) compared to areas with runoff ratio increases. Using a machine learning regression tree approach, we find that changes in canopy evaporation is the most important variable related to changes in antecedent soil water content in areas of decreased runoff ratios (with minimal changes in antecedent rainfall) while antecedent ground evaporation is the most important variable in areas of increased runoff ratios. Our results suggest that simulated interactions between vegetation greening, increasing evaporative demand, and antecedent soil drying are projected to diminish runoff associated with extreme rainfall events, with important implications for society.

soil moisture↗

Application of Machine Learning Algorithms to the Study of Noise Artifacts in Gravitational-Wave Data

The sensitivity of searches for astrophysical transients in data from the Laser Interferometer Gravitationalwave Observatory (LIGO) is generally limited by the presence of transient, non-Gaussian noise artifacts, which occur at a high-enough rate such that accidental coincidence across multiple detectors is non-negligible. Furthermore, non-Gaussian noise artifacts typically dominate over the background contributed from stationary noise. These "glitches" can easily be confused for transient gravitational-wave signals, and their robust identification and removal will help any search for astrophysical gravitational-waves. We apply Machine Learning Algorithms (MLAs) to the problem, using data from auxiliary channels within the LIGO detectors that monitor degrees of freedom unaffected by astrophysical signals. Terrestrial noise sources may manifest characteristic disturbances in these auxiliary channels, inducing non-trivial correlations with glitches in the gravitational-wave data. The number of auxiliary-channel parameters describing these disturbances may also be extremely large; high dimensionality is an area where MLAs are particularly well-suited. We demonstrate the feasibility and applicability of three very different MLAs: Artificial Neural Networks, Support Vector Machines, and Random Forests. These classifiers identify and remove a substantial fraction of the glitches present in two very different data sets: four weeks of LIGO's fourth science run and one week of LIGO's sixth science run. We observe that all three algorithms agree on which events are glitches to within 10% for the sixth science run data, and support this by showing that the different optimization criteria used by each classifier generate the same decision surface, based on a likelihood-ratio statistic. Furthermore, we find that all classifiers obtain similar limiting performance, suggesting that most of the useful information currently contained in the auxiliary channel parameters we extract is already being used. Future performance gains are thus likely to involve additional sources of information, rather than improvements in the MLAs themselves.

gravitational-wave data↗

High-Throughput Strategies that Encompass Experiments and Machine Learning to Predict the Mechanical Properties of Additive Manufactured Aerospace Alloys

Small Punch Test (SPT) uses a thin disk of material to predict mechanical properties. While SPT has existed for decades, it has been used largely as a qualitative evaluator of mechanical properties. Recent advances in computational modeling have enabled the extraction of uniaxial stress-strain response from the measured SPT load-displacement data. Due to small sample volumes and unidirectional testing, SPT is conducive to high-throughput automation and ideally suited to extract properties from high-cost materials. Aerospace alloys have been of recent interest to the Additive Manufacturing (AM) community due to AM’s unique ability to fabricate complex designs not possible, or extremely arduous, with conventional manufacturing. In this research, SPT, coupled with Materials Informatics and computational modeling, is used to develop relevant Process-Structure-Property relationships to decrease the cost and time of process optimization for AM aerospace alloys, namely Inconel 718, Inconel 625, and Niobium C103.

High-throughput Testing↗

DroughtCast: A Machine Learning Forecast of the United States Drought Monitor

Drought is one of the most ecologically and economically devastating natural phenomena affecting the United States, causing the U.S. economy billions of dollars in damage, and driving widespread degradation of ecosystem health. Many drought indices are implemented to monitor the current extent and status of drought so stakeholders such as farmers and local governments can appropriately respond. Methods toforecast drought conditions weeks to months in advance are less common but would provide a more effective early warning system to enhance drought response, mitigation, and adaptation planning. To resolve this issue, we introduce DroughtCast, a machine learning framework for forecasting the United States Drought Monitor (USDM). DroughtCast operates on the knowledge that recent anomalies in hydrology and meteorology drive future changes in drought conditions. We use simulated meteorology and satellite observed soil moisture as inputs into a recurrent neural network to accurately forecast the USDM between 1 and 12 weeks into the future. Our analysis shows that precipitation, soil moisture, and temperature are the most important input variables when forecasting future drought conditions. Additionally, a case study of the 2017 Northern Plains Flash Drought shows that DroughtCast was able to forecast a very extreme drought event up to 12 weeks before its onset. Given the favorable forecasting skill of the model, DroughtCast may provide a promising tool for land managers and local governments in preparing for and mitigating the effects of drought.

Machine Learning↗

Fine particulate concentrations over East Asia derived from aerosols measured by the Advanced Himawari Imager using machine learning

Fine particulate matter with a diameter below 2.5 μm (PM 2.5 ) is deleterious to the cardiovascular and respiratory systems. It is often difficult to assess the effects of PM 2.5 on human health over regions with limited ground monitoring sites, especially in East Asia. As an alternative, we estimated near-surface PM 2.5 concentrations by analyzing Advanced Himawari Imager (AHI) Yonsei Aerosol Retrieval (YAER) products. This study incorporates daytime data for East Asia covering the Korean Peninsula, China, Japan, Southeast Asia, and southern Mongolia. We collocated AHI YAER product pixels with meteorological, land-cover, and other ancillary data for the period from March 2018 to February 2019. To estimate PM 2.5 concentrations over wide areas spanning many countries displaying various relationships between aerosol optical depth and PM 2.5 , monthly models were developed by considering both the spatial and temporal characteristics of ground-based PM 2.5 measurements. Random forest machine learning model estimated ground-level mass concentrations of PM 2.5 ; subsequent 10-fold cross validation (CV) yielded a CV R 2 value of 0.81 and a CV root mean squared error (RMSE) of 12.3 μg m -3 . We investigated the spatial pattern of PM 2.5 concentrations over multiple countries and seasonal variation in PM 2.5 concentrations. Diurnal variation of a severe PM 2.5 event in the Korean Peninsula was investigated as a case study. The model captured the extremely heterogeneous spatial distribution of PM 2.5 concentrations peaked around local noon. To measure the capability of the developed model to estimate PM 2.5 concentrations in areas with few in-situ data, its predictive performance was evaluated using a dataset independent of the training process with an R 2 of 0.60 and RMSE of 8.18 μg m −3 . This study demonstrates the potential for satellite-based PM 2.5 estimation for areas with insufficient measuring stations.

Pm2.5↗

Communicating Metrics of Land Surface Temperature Variability Using Multi-sensor Machine Learning

Land surface temperature (LST) is a key climate observable used to detect changes in the Earth’s surface energy budget that influence carbon and water cycles. Land surface temperature exhibits strong diurnal variability, which geostationary satellites can observe at scale thanks to their temporal resolution. Due to anthropogenic climate and land use changes, the surface energy balance has been considerably modified and may be described by changes in diurnal temperature range and extremes. Using high performance computing and datasets from the NASA Earth Exchange, we exploit co-located, co-temporal observations from low-earth orbit (LEO) and geostationary (GEO) sensors to develop a deep learning-based method for LEO-to-GEO algorithm emulation. Our model is trained to predict MODIS Terra LST from GOES-16 thermal bands and achieves validation error <2K. Application of the model to unseen times of day (observed by MODIS Aqua) and a new GEO sensor (Himawari-8) observing an unseen spatial domain, demonstrate the generalization of the deep learning model across space, time and spectra. Further, time series clustering approaches are examined with the objective of identifying key indicators of change in diurnal cycling and extremes on a continental scale. Communicating LST variability observed by geostationary satellites can have impacts in multiple disciplines, from understanding of snow, vegetation and soil dynamics, to recognizing trends in heat events relevant to human health.

Kate Duffy↗

In-Situ Scanning Electron Microscope Experiments for Microscale Mechanical Testing and Validated Modeling of Fiber Reinforced Thermoplastics

A novel, in-situ, scanning electron microscope (SEM) mechanical testing capability for materials at the microscale which provides experimental validation to a machine learning (ML) toolset for full-field validation of physics-based micromechanics models is being developed by researchers at NASA Glenn Research Center. These are enabling technologies for the integration of multiscale digital twins for materials into system level models which will result in the improved performance, material discovery, reduced production cost and time, rapid characterization, and prognostic structural health monitoring (SHM) for materials and structures for extreme environments in support of NASA space exploration missions. In order to bridge the material structure-to-system gap for digital twins, physics-based models must be experimentally validated at multiple length scales. Seminal microscale experiments, conducted at the Air Force Research Laboratory (AFRL), were limited to transverse compression of single-layer, unidirectional thermoset polymer matrix composite (PMC) micropillar specimens [1]. The early phases of the current project followed those initial results and setup to reproduce the compression testing of PMC material on the custom-built piezoelectric actuated micromechanical testing rig built by MicroTesting Solutions LLC. In this work, samples of thermoplastic PMC material were first machined into 3 mm cubes, and then further machining and final milling was done using a Focused Ion Beam (FIB). The initial experiment was done on a pillar roughly 20 µm x 20 µm x 40 µm tall. Additional pillars were milled with final sizes ranging from 20 µm x 20 µm x 40 µm tall to 40 µm x 40 µm x 65 µm tall. A speckle pattern for in-situ full-field measurements using Digital Image Correlation (DIC) was applied with platinum, which was coated on the surface, and then the FIB was used to mill away some of the coating to produce an irregular pattern of Pt on the pillar surface. The samples were loaded into the custom testing rig and placed into the SEM and loaded under compression until failure. Images were collected in the SEM during testing. Post-processing of the images was conducted using DIC to obtain full-field displacement and strain measurements elucidating the role of the matrix as well as fiber-fiber interaction at the microscale within the composite subjected to compression loading well into the non-linear regime of the material. Moreover, the evolution of fiber-matrix debonding and matrix cracking is observed in-situ at the microscale. This data, along with images segmented with a newly developed ML toolset [2], was used to create and validate physics-based micromechanics models. An image of the failed micropillar is shown in Figure 1. The techniques developed in the initial compression experiment was tailored to the validation needs of the models and expanded to include different sized samples as well as possibly tension and fatigue.

Laura Wilson↗

Decoding the effects of synonymous variants

Synonymous single nucleotide variants (sSNVs) are common in the human genome but are often overlooked. However, sSNVs can have significant biological impact and may lead to disease. Existing computational methods for evaluating the effect of sSNVs suffer from the lack of gold-standard training/evaluation data and exhibit over-reliance on sequence conservation signals. We developed synVep (synonymous Variant effect predictor), a machine learning-based method that overcomes both of these limitations. Our training data was a combination of variants reported by gnomAD (observed) and those unreported, but possible in the human genome (generated). We used positive-unlabeled learning to purify the generated variant set of any likely unobservable variants. We then trained two sequential extreme gradient boosting models to identify subsets of the remaining variants putatively enriched and depleted in effect. Our method attained 90% precision/recall on a previously unseen set of variants. Furthermore, although synVep does not explicitly use conservation, its scores correlated with evolutionary distances between orthologs in cross-species variation analysis. synVep was also able to differentiate pathogenic vs. benign variants, as well as splice-site disrupting variants (SDV) vs. non-SDVs. Thus, synVep provides an important improvement in annotation of sSNVs, allowing users to focus on variants that most likely harbor effects.

Zishuo Zeng↗