Comparative modeling reveals the molecular determinants of aneuploidy fitness cost in a wild yeast model
Code repository for aneuploidy model and gene machine learning model
SEARCH · Engineering Papers
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Code repository for aneuploidy model and gene machine learning model
As the incidence of obesity and associated negative health consequences is rising, it becomes crucial to monitor the dietary choices of individuals. Unfortunately, traditional methods to collect this information involve collecting food frequency questionnaires from individuals using paper. Electronic food trackers have been developed to collect food data, but they require participants to manually label and describe the content of their meals, and which may be difficult for researchers to interpret in a standardized fashion. Machine learning, however, provides an easy and efficient method for both participants and researchers to label food items with standardized descriptions. This project aims to create a prototype phone application that can identify and label photos of apples. This is done by making a machine learning model through Turicreate, a python module, which is then implemented into an iOS app through Xcode and Swift. The modules used in Swift include CoreML and AVFoundation. This machine learning application will be incorporated with a MealLogger phone app that is also under development. The MealLogger app will be used to keep track of participants' calorie intake and other personal details throughout the sleep study. The machine learning model will present several potential identities of the foods found in the photo, and the user will only need to select the correct option. This will be a user-friendly method for participants to easily log their food consumption without the hard work of manually inputting each and every description. Some limitations to this project include the wide variety of food, including those within different cultures. To deal with this, the model will include the most generic food categories, which the participant may select, and produce a drop-down menu of more specific dishes under that specified category, with the option of self-input. Additional questionnaires may be implemented according to the food type selected This will allow the process to be quick and easy, but also specific for the purpose of analysis. The release of the application will require a much longer process, but the machine learning prototype presents a first step toward an application that may change data analysis for researchers interested in collecting food intake from individuals living in the real world.
Background: Checking the connectivity (structure) of complex Metabolic Reaction Networks(MRNs) models proposed for new microorganisms with promising properties is an importantgoal for chemical biology. Objective: In principle, we can perform a hand-on checking (Manual Curation). However, this is achallenging task due to the high number of combinations of pairs of nodes (possible metabolic reactions). Results: The CPTML linear model obtained using the LDA algorithm is able to discriminate nodes(metabolites) with the correct assignation of reactions from incorrect nodes with values of accuracy,specificity, and sensitivity in the range of 85-100% in both training and external validation dataseries. Methods: In this work, we used Combinatorial Perturbation Theory and Machine Learning techniquesto seek a CPTML model for MRNs >40 organisms compiled by Barabasis’ group. First, wequantified the local structure of a very large set of nodes in each MRN using a new class of node indexcalled Markov linear indices fk. Next, we calculated CPT operators for 150000 combinationsof query and reference nodes of MRNs. Last, we used these CPT operators as inputs of differentML algorithms. Conclusion: Meanwhile, PTML models based on Bayesian network, J48-Decision Tree and RandomForest algorithms were identified as the three best non-linear models with accuracy greaterthan 97.5%. The present work opens the door to the study of MRNs of multiple organisms usingPTML models.
Due to the growing amount of data from in-situ sensors in environmental monitoring, it becomes necessary to automatically detect anomalous data points. Nowadays, this is mainly performed using supervised machine learning models, which need a fully labelled data set for their training process. However, the process of labelling data is typically cumbersome and, as a result, a hindrance to the adoption of machine learning methods for automated anomaly detection. In this work, we propose to address this challenge by means of active learning. This method consists of querying the domain expert for the labels of only a selected subset of the full data set. We show that this reduces the time and costs associated to labelling while delivering the same or similar anomaly detection performances. Finally, we also show that machine learning models providing a nonlinear classification boundary are to be recommended for anomaly detection in complex environmental data sets.
This data package is associated with the publication “Combined effects of stream hydrology and land use on basin‐scale hyporheic zone denitrification in the Columbia River Basin”, published in Water Resource Research (Son et al.2022) available at https://doi.org/10.1029/2021WR031131. This data package includes the key model inputs/outputs of the river corridor model for the Columbia River Basin (CRB) and the model source codes used in the manuscript. The model is a carbon-nitrogen-coupled river corridor model (RCM), and the model is used to quantify hyporheic zone (HZ) denitrification at the NHDPLUS stream reach scales. The RCM used in this study combines empirical substrate models derived from observations and three microbially driven reactions, including two-step denitrification and aerobic respiration, are considered within the HZ. The key input data of the model are exchange flux, residence time, and stream solute (dissolved organic carbon (DOC), dissolved oxygen (DO), and nitrate concentrations). These inputs are constant over time and represent long-term averaged values. This study uses the RCM to explore the spatial patterns of HZ denitrification across reaches with different sizes and land use in the CRB. Our main objective is to use the RCM as a virtual reality model, and the machine-learning models as surrogates that encapsulate the complexities of the physics-based model while identifying the importance of different variables that are not evident in the model conceptualization. We do not include a direct comparison of the modeled HZ denitrification and measurements; however, the RCM can capture the overall spatial patterns of the HZ denitrification because the model inputs and its reaction networks are based on well-established theory and a physical-based model. The combination of the model-based predictions and a machine-learning approach (e.g., random forest) is used to improve our understanding of what variables of the model are associated with spatial patterns of the modeled denitrification across reaches with different sizes and land uses, and to develop a proxy model using measurable variables to reproduce the simulated patterns.This dataset contains five folders: (1) model_inputs, (2) model_outputs, (3) Rscripts, (4) figures, and (5) model_codes. It also contains a readme, file level metadata (FLMD), and data dictionary (dd). Please see the FLMD for a list of all the files contained in this data package and descriptions for each. The model_inputs folder contains the model inputs used to drive the model simulations. The model_outputs folder contains key model output files from the river corridor model. The Rscripts folder contains the Rscripts for pre- and post- processing model results. The figures folder contains the raw figures associated with the manuscript. The model_codes folder includes key model source codes/input files. All files are .jpg, .jpeg, .out, .e, .od, .dat, .sub, .F90, .0, .R, .sbx, .cpg, .sbn, .shx, .shp, .dbf, .prj, .tfw, .tif, .xml, .pdf, or .csv.
Numerous phenomenological nuclear models have been proposed to describe specific observables within different regions of the nuclear chart. However, developing a unified model that describes the complex behavior of all nuclei remains an open challenge. Here, we explore whether symbolic Machine Learning (ML) can rediscover traditional nuclear physics models or identify alternatives with improved simplicity, fidelity, and predictive power. To address this challenge, we developed a Multi-objective Iterated Symbolic Regression approach that handles symbolic regressions over multiple target observables, accounts for experimental uncertainties and is robust against high-dimensional problems. As a proof of principle, we applied this method to describe the nuclear binding energies and charge radii of light and medium mass nuclei. Our approach identified simple analytical relationships based on the number of protons and neutrons, providing interpretable models with precision comparable to state-of-the-art nuclear models. Additionally, we integrated this ML-discovered model with an existing complementary model to estimate the limits of nuclear stability. These results highlight the potential of symbolic ML to develop accurate nuclear models and guide our description of complex many-body problems.
Accurate estimates of the incidence of infectious diseases are key for the control of epidemics. However, healthcare systems are often unable to test the population exhaustively, especially when asymptomatic and paucisymptomatic cases are widespread; this leads to significant and systematic under-reporting of the real incidence. Here, we propose a machine learning approach to estimate the incidence of a pandemic in real-time, using reported cases and the overall test rate. In particular, we use Bayesian symbolic regression to automatically learn the closed-form mathematical models that most parsimoniously describe incidence. We develop and validate our models using COVID-19 incidence values for nine different countries, confirming their ability to accurately predict daily incidence. Remarkably, despite the differences in epidemic trajectories and dynamics across countries, we find that a single model for all countries offers a more parsimonious description and is more predictive of actual incidence compared to separate models for each country. Our results show the potential to accurately model incidence in real-time using closed-form mathematical models, providing a valuable tool for public health decision-makers.
Phase 1 (Original CRADA, plus no-cost extension modifications #1-3, 6/1/2017 to 3/13/2021): The Australian Department of Defence (AUDoD) is performing accelerated aging tests of Li-ion batteries to benchmark their reliability and degradation characteristics. Using its previously developed battery lifetime predictive model framework, the National Laboratory of the Rockies (NLR) will develop analytical models based the AUDoD data to predict lifetime of the multiple Li-ion battery chemistries under real-world use scenarios of interest to AUDoD. The NLR model is based on physical degradation mechanisms encountered by Li-ion batteries and has been previously validated. Phase 2 (CRADA modification #4, plus no-cost extension modification #5, 2/22/2021 to 3/30/2025): Train and support Australian Department of Defence personnel to use NLR software for model-based estimation of Li-ion battery lifetime using accelerated battery aging data collected by the Australian Department of Defence. Under separate DOE funding from 2019 to 2021, NLR enhanced its battery life-prediction software using machine learning algorithms to automate portions of the model-fitting process, requiring significantly less labor and expert judgment and also adding uncertainty quantification, increasing statistical rigor. Under Phase 2, NLR will customize NLR Software and provide it to AuDoD. NLR will enhance its NLR Model to capture aging modes of AuDoD's multi-cell modules, including cell-balancing effects. NLR will develop example single-cell and multi-cell models based on one AuDoD battery aging dataset. NLR will train AuDoD personnel on NLR Software. By the conclusion of the project, NLR will have provided AuDoD the training materials, a user manual and software needed to perform their own analysis of additional and/or future battery aging datasets.
The parameterization of key photosynthesis parameters is one of the key uncertain sources in modeling ecosystem gross primary productivity (GPP). Solar-induced chlorophyll fluorescence (SIF) offers a good proxy for GPP since it marks the actual process of photosynthesis; while machine learning (ML) provides a robust approach to model the GPP-SIF relationship. Here, we trained the boosted regressing tree (BRT) and the Random Forest ML models with Greenhouse Gases Observing Satellite SIF data and in situ GPP observations from 49 eddy covariance towers. These trained ML GPP-SIF models were fed into the Energy Exascale Earth System Model (E3SM) Land Model (ELM) to generate ELM-simulated global SIF estimates, which were then benchmarked against satellite SIF observations with a surrogate modeling approach. Our results indicated good modeling performance of the ML-based GPP-SIF relationship. The ELM model when fed with the ML GPP-SIF models also can well predict the spatial-temporal variations in SIF. We also found high model accuracy for the surrogate modeling. Model parameter sensitivity analysis suggested that the fraction of leaf nitrogen in RuBisCO (flnr) is the most sensitive parameter to the SIF; other sensitive parameters include the Ball-Berry stomatal conductance slope (mbbopt) and the vcmax entropy (vcmaxse). The posterior uncertainty in simulated GPP was greatly reduced after benchmarking, and the model produced improved spatial patterns of mean GPP relative to FLUXCOM GPP. Our integrated approach provides a new avenue for improving land models and using remote-sensing SIF, which can be further improved in the future with more ground- and satellite-based observations.
Here, in this work we present deep neural network regression machine learning models (ML) for predicting the average voltage and the percentage change in volume of battery electrodes upon charging and discharging with metal ions. Our models exhibit good performance as measured by the average mean absolute error obtained from a 10-fold cross-validation as well as on independent test sets. We further assess the robustness our ML models by investigating their screening potential beyond the training database. We produce novel Na-ion electrodes by systematically replacing Li-ions in the original database by Na-ions, and then selecting a set of 22 electrodes that exhibit a good performance in energy density as well as small volume variations upon charging and discharging, as predicted by the machine learning model. The ML predictions for these new materials are then compared to quantum-mechanics based calculations. Our results reaffirm the significant role of machine learning techniques in the exploration of materials for battery applications.
The ability to rapidly screen material performance in the vast space of high entropy alloys is of critical importance to efficiently identify optimal hydride candidates for various use cases. Given the prohibitive complexity of first principles simulations and large-scale sampling required to rigorously predict hydrogen equilibrium in these systems, we turn to compositional machine learning models as the most feasible approach to screen on the order of tens of thousands of candidate equimolar high entropy alloys (HEAs). Critically, we show that machine learning models can predict hydride thermodynamics and capacities with reasonable accuracy (e.g. a mean absolute error in desorption enthalpy prediction of ~5 kJ mol H 2 –1 ) and that explainability analyses capture the competing trade-offs that arise from feature interdependence. We can therefore elucidate the multi-dimensional Pareto optimal set of materials, i.e., where two or more competing objective properties can't be simultaneously improved by another material. This provides rapid and efficient down-selection of the highest priority candidates for more time-consuming density functional theory investigations and experimental validation. Various targets were selected from the predicted Pareto front (with saturation capacities approaching two hydrogen per metal and desorption enthalpy less than 60 kJ mol H 2 –1 ) and were experimentally synthesized, characterized, and tested amongst an international collaboration group to validate the proposed novel hydrides. Finally, additional top-predicted candidates are suggested to the community for future synthesis efforts, and we conclude with an outlook on improving the current approach for the next generation of computational HEA hydride discovery efforts.
HLS4ml (high level synthesis for machine learning) Is a Python package used to translate commonly used open-source machine learning models into HLS. This is useful in machine learning applications on FPGAs. Machine learning algorithms are only as fast as the hardware that they are used on, and some applications require high speed without sacrificing accuracy. In these situations, an FPGA is a good choice since it is faster than a CPU or a GPU, but programming an FPGA is difficult. This is where HLS4ml can be used to simplify the process, as a well-known learning model can be converted to HLS and more easily deployed onto an FPGA. There are many use cases for a machine learning algorithm running on an FPGA. For example, detectors in a particle accelerator cannot keep every event that they detect, and so a computer must decide which events to keep and which to discard. Using an FPGA with a machine learning algorithm would be a good way to keep as many events as possible.
hls4ml (high level synthesis for machine learning) Is a Python package used to translate commonly used open-source machine learning models into HLS. This is useful in machine learning applications on FPGAs. Machine learning algorithms are only as fast as the hardware that they are used on, and some applications require high speed without sacrificing accuracy. In these situations, an FPGA is a good choice since it is faster than a CPU or a GPU, but programming an FPGA is difficult. This is where hls4ml can be used to simplify the process, as a well-known learning model can be converted to HLS and more easily deployed onto an FPGA. There are many use cases for a machine learning algorithm running on an FPGA. For example, detectors in a particle accelerator cannot keep every event that they detect, and so a computer must decide which events to keep and which to discard. Using an FPGA with a machine learning algorithm would be a good way to keep as many events as possible.
Hybrid composites combine two or more different fillers to achieve multifunctional or advanced material properties, such as lightweight and enhanced mechanical properties. The properties of the composites significantly depend on their microstructures, which can be tailored via advanced 3D printing processes. Understanding the process-structure-property relationships is critical to enable the design and engineering of novel hybrid composites for applications in aerospace, automotive, and protective coatings. Here, for this work, we develop 3D printable and lightweight hybrid composites and leverage the conventional design of experiments, a theoretical hybrid model, and an image-driven machine learning (ML) method to investigate their mechanical behaviors. The hybrid composites are formulated with elastomer matrix, microfillers, and thin-shell particles, enabling a significant degree of design freedom of microstructures with densities and mechanical properties varying up to 70% and 91%, respectively. Our statistical analysis indicates that the 3D printing path direction and the microfibers fraction are dominating process parameters with contribution percentages of 45.3% and 57.7% on the specific stiffness and strength, respectively. A hybrid mechanics model is developed based on a simple Weibull distribution function and classical single-filler models to effectively capture the variations in mechanical properties, however, it overestimates the values due to its statistical constraints and idealization of experimental uncertainty. The image-driven ML model leverages the microscale images directly without losing the structural details, shows more accurate predictions with experimental data, and has 48.6% lower root mean square error than the theoretical model.
There is growing support and interest in postsecondary interdisciplinary environmental education which integrate concepts and disciplines in addition to providing varied perspectives. There is a need to assess student learning in these programs as well as rigorous evaluation of educational practices, especially of complex synthesis concepts. This work tests a text classification machine learning model as a tool to assess student systems thinking capabilities using two questions anchored by the Food-Energy-Water (FEW) Nexus phenomena by answering two questions (1) Can machine learning models be used to identify instructor-determined important concepts in student responses? (2) What do college students know about the interconnections between food, energy and water, and how have students assimilated systems thinking into their constructed responses about FEW? Reported here are a broad range of model performances across 26 text classification models associated with two different assessment items, with model accuracy ranging from 0.755 to 0.992. Expert-like responses were infrequent in our dataset compared to responses providing simpler, incomplete explanations of the systems presented in the question. For those students moving from describing individual effects to multiple effects, their reasoning about the mechanism behind the system indicates advanced systems thinking ability. Specifically, students exhibit higher expertise for explaining changing water usage than discussing tradeoffs for such changing usage. This research represents one of the first attempts to assess the links between foundational, discipline-specific concepts and systems thinking ability. These text classification approaches to scoring student FEW Nexus Constructed Responses (CR) indicate how these approaches can be used, in addition to several future research priorities for interdisciplinary, practice-based education research. Development of further complex question items using machine learning would allow evaluation of the relationship between foundational concept understanding and integration of those concepts as well as more nuanced understanding of student comprehension of complex interdisciplinary concepts.
Introduction The field of machine learning and its subfield of deep learning have grown rapidly in recent years. With the speed of advancement, it is nearly impossible for data scientists to maintain expert knowledge of cutting-edge techniques. This study applies human factors methods to the field of machine learning to address these difficulties. Methods Using semi-structured interviews with data scientists at a National Laboratory, we sought to understand the process used when working with machine learning models, the challenges encountered, and the ways that human factors might contribute to addressing those challenges. Results Results of the interviews were analyzed to create a generalization of the process of working with machine learning models. Issues encountered during each process step are described. Discussion Recommendations and areas for collaboration between data scientists and human factors experts are provided, with the goal of creating better tools, knowledge, and guidance for machine learning scientists.
Air pollution in the Hindu Kush Himalayan (HKH) region of South Asia is a severe issue, as increases in emissions over the past two decades have degraded air quality (AQ) across the region, which poses major threats to human health, the ecosystem, climate, and agriculture. A diversity of anthropogenic and natural emission sources including transportation, power plants, industries, open biomass burning of crop residue, forest fires, cooking and heating fires, and dust storms contribute to unhealthy AQ and transboundary pollution issues in the region. Further complicating matters is the importance of meteorology and terrain on AQ, especially in the Kathmandu Valley where extreme haze episodes frequently develop from the atmospherically stable weather conditions during the winter monsoon. This study uses state-of-the-art satellite observations and modeling capabilities in conjunction with machine learning techniques to develop a comprehensive toolkit for enhancing AQ monitoring and forecasting in HKH. The toolkit incorporates new generation satellite observations from the TROPOspheric Monitoring Instrument (TROPOMI), Geostationary Environment Monitoring Spectrometer (GEMS), and Advanced Meteorological Imager (AMI), which provide unprecedented resolution on aerosols and trace gases, including nitrogen dioxide, formaldehyde, sulfur dioxide, carbon monoxide, and ozone, and aerosol optical depth. Value-added products, such as level 4 PM2.5 products, are developed from the suite of satellite observations to further improve AQ monitoring capabilities in the region. The satellite products are also used to assimilate a high-resolution chemical transport model tailored for the HKH region, which is providing daily, 54-hour AQ forecasts with horizontal grid spacings of 12- and 4-km. This presentation will provide an overview of the suite of satellite- and model-based products in the AQ toolkit and application and performance of the toolkit for AQ monitoring and forecasting in HKH.
Air pollution in the Hindu Kush Himalayan (HKH) region of South Asia is a severe issue, as increases in emissions over the past two decades have degraded air quality (AQ) across the region, which poses major threats to human health, the ecosystem, climate, and agriculture. A diversity of anthropogenic and natural emission sources including transportation, power plants, industries, open biomass burning of crop residue, forest fires, cooking and heating fires, and dust storms contribute to unhealthy AQ and transboundary pollution issues in the region. Further complicating matters is the importance of meteorology and terrain on AQ, especially in the Kathmandu Valley where extreme haze episodes frequently develop from the atmospherically stable weather conditions during the winter monsoon. This study uses state-of-the-art satellite observations and modeling capabilities in conjunction with machine learning techniques to develop a comprehensive toolkit for enhancing AQ monitoring and forecasting in HKH. The toolkit incorporates new generation satellite observations from the TROPOspheric Monitoring Instrument (TROPOMI), Geostationary Environment Monitoring Spectrometer (GEMS), and Advanced Meteorological Imager (AMI), which provide unprecedented resolution on aerosols and trace gases, including nitrogen dioxide (NO 2 ), formaldehyde (CH2O), sulfur dioxide (SO 2 ), carbon monoxide (CO), and ozone (O 3 ), and aerosol optical depth (AOD). Value-added products [e.g., Particulate matter with diameters less than 2.5 micrometers (PM2.5)] are developed from the suite of satellite observations to further improve AQ monitoring capabilities in the region. The satellite products are also used to assimilate a high-resolution chemical transport model tailored for the HKH region, which is providing daily, 54-hour AQ forecasts with horizontal grid spacings of 12- and 4-km. This presentation will provide an overview of the suite of satellite- and model-based products in the AQ toolkit and application and performance of the toolkit for AQ monitoring and forecasting in HKH.