Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “annotations”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 469 records · Page 26

Increased inflammation as well as decreased endoplasmic reticulum stress and translation differentiate pancreatic islets from donors with pre-symptomatic stage 1 type 1 diabetes and non-diabetic donors

Aims/hypothesis Progression to type 1 diabetes is associated with genetic factors, the presence of autoantibodies and a decline in beta cell insulin secretion in response to glucose. Very little is known regarding the molecular changes that occur in human insulin-secreting beta cells prior to the onset of type 1 diabetes. Herein, we applied an unbiased proteomics approach to identify changes in proteins and potential mechanisms of islet dysfunction in islet-autoantibody-positive organ donors with pre-symptomatic stage 1 type 1 diabetes (HbA1c ≤42 mmol/mol [6.0%]). We aimed to identify pathways in islets that are indicative of beta cell dysfunction. Methods Multiple islet sections were collected through laser microdissection of frozen pancreatic tissues from organ donors positive for single or multiple islet autoantibodies (AAb + , n=5), and age (±2 years)- and sex-matched non-diabetic (ND) control donors (n=5) obtained from the Network for Pancreatic Organ donors with Diabetes (nPOD). Islet sections were subjected to MS-based proteomics and analysed with label-free quantification followed by pathway and functional annotations. Results Analyses resulted in ~4500 proteins identified with low false discovery rate (<1%), with 2165 proteins reliably quantified in every islet sample. We observed large inter-donor variations that presented a challenge for statistical analysis of proteome changes between donor groups. We therefore focused on only the donors with stage 1 type 1 diabetes who were positive for multiple autoantibodies (mAAb + , n=3) and genetic risk compared with their matched ND controls (n=3) for the final statistical analysis. Approximately 10% of the proteins (n=202) were significantly different (unadjusted p<0.025, q<0.15) for mAAb + vs ND donor islets. The significant alterations clustered around major functions for upregulation in the immune response and glycolysis, and downregulation in endoplasmic reticulum (ER) stress response as well as protein translation and synthesis. The observed proteome changes were further supported by several independent published datasets, including a proteomics dataset from in vitro proinflammatory cytokine-treated human islets and single-cell RNA-seq datasets from AAb + individuals. Conclusions/interpretation In situ human islet proteome alterations in stage 1 type 1 diabetes centred around several major functional categories, including an expected increase in immune response genes (elevated antigen presentation/HLA), with decreases in protein synthesis and ER stress response, as well as compensatory metabolic response. The dataset serves as a proteomics resource for future studies on beta cell changes during type 1 diabetes progression and pathogenesis. Data availability The LC-MS raw datasets that support the findings of this study have been deposited in the online repository: MassIVE (https://massive.ucsd.edu/ProteoSAFe/static/massive.jsp) with accession no. MSV000090212.

Autoantibody-positive↗

Characterization of the biofilm landscape of Bacillus subtilis by spatial microproteomics

Bulk proteomics has been demonstrated to differentiate subpopulations within bacterial colonies, yet advanced analyses by mass spectrometry imaging (MSI) hold even greater promise for the future. This technology can enable high-throughput spatial phenotyping that can reshape biological discovery by providing visualization of components of various biomolecular mechanisms. With high mass resolving power and high spatial resolution analyses being routine, we can confidently enable intact protein imaging directly from samples with minimal preparation. Pairing those analyses with bulk experimental libraries can provide high confidence in annotations of post-translational modifications (PTMs) and truncations. Revealing PTM localization within the samples unlocks a direct window into unknown biology at the microscale. However, top-down proteomics (TDP) is not commonplace for microbial species, largely due to challenges in identifying detected peptides and proteins; considering the theoretical proteome of even the well-studied model bacterium Bacillus subtilis was only partially mapped recently. With little still known about the form and function of many of these proteins – let alone proteoforms, where PTMs and truncations of the same protein may possess unique physiological roles – there is a wealth of work to be done. Here we jointly apply TDP and MSI to describe the microscale spatial proteomic landscape within B. subtilis and further demonstrate the feasibility of detecting differentiated subpopulations through proteoforms across the biofilm landscape.

bacterial biofilms↗

Systems analysis of Lipomyces starkeyi during growth on various plant-based sugars

Oleaginous yeasts have received significant attention due to their substantial lipid storage capability. The accumulated lipids can be utilized directly or processed into various bioproducts and biofuels. Lipomyces starkeyi is an oleaginous yeast capable of using multiple plant-based sugars, such as glucose, xylose, and cellobiose. It is, however, a relatively unexplored yeast due to limited knowledge about its physiology. In this study, we have evaluated the growth of L. starkeyi on different sugars and performed transcriptomic and metabolomic analyses to understand the underlying mechanisms of sugar metabolism. Principal component analysis showed clear differences resulting from growth on different sugars. We have further reported various metabolic pathways activated during growth on these sugars. We also observed non-specific regulation in L. starkeyi and have updated the gene annotations for the NRRL Y-11557 strain. Furthermore, this analysis provides a foundation for understanding the metabolism of these plant-based sugars and potentially valuable information to guide the metabolic engineering of L. starkeyi to produce bioproducts and biofuels.

59 BASIC BIOLOGICAL SCIENCES↗

The Ontology of Biological Attributes (OBA)—computational traits for the life sciences

Abstract Existing phenotype ontologies were originally developed to represent phenotypes that manifest as a character state in relation to a wild-type or other reference. However, these do not include the phenotypic trait or attribute categories required for the annotation of genome-wide association studies (GWAS), Quantitative Trait Loci (QTL) mappings or any population-focussed measurable trait data. The integration of trait and biological attribute information with an ever increasing body of chemical, environmental and biological data greatly facilitates computational analyses and it is also highly relevant to biomedical and clinical applications. The Ontology of Biological Attributes (OBA) is a formalised, species-independent collection of interoperable phenotypic trait categories that is intended to fulfil a data integration role. OBA is a standardised representational framework for observable attributes that are characteristics of biological entities, organisms, or parts of organisms. OBA has a modular design which provides several benefits for users and data integrators, including an automated and meaningful classification of trait terms computed on the basis of logical inferences drawn from domain-specific ontologies for cells, anatomical and other relevant entities. The logical axioms in OBA also provide a previously missing bridge that can computationally link Mendelian phenotypes with GWAS and quantitative traits. The term components in OBA provide semantic links and enable knowledge and data integration across specialised research community boundaries, thereby breaking silos.

59 BASIC BIOLOGICAL SCIENCES↗

Drought shifts dissolved organic matter sources from above- to belowground and stress-induced processes in Amazon white-sand forests

White-sand forests contribute significantly to dissolved organic matter (DOM) production in the central Amazon, forming blackwater rivers that dominate organic matter export from the Amazon basin to the ocean. Despite their importance in controlling DOM export, white-sand forests are understudied, and it remains unclear whether systematic changes in the formation of blackwater DOM occur and how seasonal variations and extremes like El Niño-associated droughts impact them. We collected soil porewater from two central Amazon white-sand forests for two years, spanning a wet La Niña year followed by an El Niño drought year. The molecular composition of DOM was analyzed using high-resolution mass spectrometry, and correlation network analysis was employed to identify ecologically meaningful DOM subsets. Using additional chemical characterization, database annotations, correlation with 14C-age of DOM and climatic variables, and ecological null modeling, we propose five distinct DOM sources: plant litter and throughfall, soil organic matter (SOM) decomposition, root exudation, and two drought response subsets of likely microbial and plant origin. During drought conditions, aboveground plant-derived compounds decreased, while SOM products, root exudates, and drought response compounds increased. These drought responses were qualitatively similar in both years but notably amplified in the drier El Niño year. Drought amplified deterministic control over DOM composition, indicating that DOM reflected directed biological responses and that future droughts are likely to generate similar shifts. Overall, drought substantially altered belowground carbon cycling by shifting DOM sources and inducing stress responses, effects expected to recur and potentially intensify under future climate scenarios.

Lange, Dan F.↗

Optimizing inference of segmentation on high-resolution images in MLExchange

MLExchange is a machine learning (ML) operations platform providing web user-interfaces (UIs) for data visualization and analysis pipelines at synchrotron facilities. Among these UIs is the segmentation app which helps synchrotron users utilize ML algorithms to automatically segment high-resolution scientific images with minimal manual annotation effort. In this work, we share code optimizations that significantly speed up the segmentation inference workflow of large data in short time. By optimizing the sequence of CPU-GPU data transfers and introducing CPU parallelization to key operations, we improve the per-device, per-image frame computational efficiency and observe close to 3×$$\times$$ speedup over the original segmentation inference workflow run time when utilizing a single GPU. Further adaptations enabling multi-GPU inference yield more than 40×$$\times$$ speedup with 100 GPUs compared to the optimized single GPU inference workflow. This acceleration of the segmentation inference workflow will provide MLExchange users with easy access to segmentation results with little wait time.

Lu, Shizhao↗

Interpreting the Lipidome: Bioinformatic Approaches to Embrace the Complexity

Background Improvements in mass spectrometry (MS) technologies coupled with bioinformatics developments have allowed considerable advancement in the measurement and interpretation of lipidomics data in recent years. Since research areas employing lipidomics are rapidly increasing, there is a great need for bioinformatic tools that capture and utilize the complexity of the data. Currently, the diversity and complexity within the lipidome is often concealed by summing over or averaging individual lipids up to (sub)class-based descriptors, losing valuable information about biological function and interactions with other distinct lipids molecules, proteins and/or metabolites. Aim of review To address this gap in knowledge, novel bioinformatics methods are needed to improve identification, quantification, integration and interpretation of lipidomics data. The purpose of this mini-review is to summarize exemplary methods to explore the complexity of the lipidome. Key scientific concepts of review Here we describe six approaches that capture three core focus areas for lipidomics: (1) lipidome annotation including a resolvable database identifier, (2) interpretation via pathway- and enrichment-based methods, and (3) understanding complex interactions to emphasize specific steps in the analytical process and highlight challenges in analyses associated with the complexity of lipidome data.

Kyle, Jennifer E.↗

Multiscale Modeling Meets Machine Learning: What Can We Learn?

Machine learning is increasingly recognized as a promising technology in the biological, biomedical, and behavioral sciences. There can be no argument that this technique is incredibly successful in image recognition with immediate applications in diagnostics including electrophysiology, radiology, or pathology, where we have access to massive amounts of annotated data. However, machine learning often performs poorly in prognosis, especially when dealing with sparse data. This is a field where classical physics-based simulation seems to remain irreplaceable. In this review, we identify areas in the biomedical sciences where machine learning and multiscale modeling can mutually benefit from one another: Machine learning can integrate physics-based knowledge in the form of governing equations, boundary conditions, or constraints to manage ill-posted problems and robustly handle sparse and noisy data; multiscale modeling can integrate machine learn- ing to create surrogate models, identify system dynamics and parameters, analyze sensitivities, and quantify uncertainty to bridge the scales and understand the emergence of function. With a view towards applications in the life sciences, we discuss the state of the art of combining machine learning and multiscale modeling, identify applications and opportunities, raise open questions, and address potential challenges and limitations. We anticipate that it will stimulate discussion within the community of computational mechanics and reach out to other disciplines including mathematics, statistics, computer science, artificial intelligence, biomedicine, systems biology, and precision medicine to join forces towards creating robust and efficient models for biological systems.

machine learning, multiscale modeling, physics-bas↗

An Open Combinatorial Diffraction Dataset Including Consensus Human and Machine Learning Labels with Quantified Uncertainty for Training New Machine Learning Models

Modern machine learning and autonomous experimentation schemes in materials science rely on accurate analysis of the data ingested by these models. Unfortunately, accurate analysis of the underlying data can be difficult, even for domain experts, complicating the training of the models intended to drive experiments. This is especially true when the goal is to identify the presence of weak signatures in diffraction or spectroscopic datasets. In this work, we examine a set of as-obtained diffraction data that track the phase transition from monoclinic to tetragonal in a Nb-doped VO2 film as a function of temperature and dopant concentration. We then task a set of domain experts and a set of machine learning experts with identifying which phase is present in each diffraction pattern manually and algorithmically, respectively; in both cases, the labels can vary dramatically, especially at the phase boundaries. We use the mode of the labels and the Shannon entropy as a method to capture, preserve and propagate consensus labels and their variance. Further we use the expert labels as a benchmark and demonstrate the use of Shannon entropy weighted scoring to test the performance of machine learning generated labels. Finally, we propose a material data challenge centered around generating improved labeling algorithms. This real-world dataset curated with expert labels can act as test bed for new algorithms. The raw data, annotations and code used in this study are all available online at data.gov and the interested reader is encouraged to replicate and improve the existing models

97 MATHEMATICS AND COMPUTING↗

Automated Grain Boundary (GB) Segmentation and Microstructural Analysis in 347H Stainless Steel Using Deep Learning and Multimodal Microscopy

Austenitic 347H stainless steel offers superior mechanical properties and corrosion resistance required for extreme operating conditions such as high temperature. The change in microstructure due to composition and process variations is expected to impact material properties. Identifying microstructural features such as grain boundaries thus becomes an important task in the process-microstructure-properties loop. Applying convolutional neural network (CNN)-based deep learning models is a powerful technique to detect features from material micrographs in an automated manner. In contrast to microstructural classification, supervised CNN models for segmentation tasks require pixel-wise annotation labels. However, manual labeling of the images for the segmentation task poses a major bottleneck for generating training data and labels in a reliable and reproducible way within a reasonable timeframe. Microstructural characterization especially needs to be expedited for faster material discovery by changing alloy compositions. Here, in this study, we attempt to overcome such limitations by utilizing multimodal microscopy to generate labels directly instead of manual labeling. We combine scanning electron microscopy images of 347H stainless steel as training data and electron backscatter diffraction micrographs as pixel-wise labels for grain boundary detection as a semantic segmentation task. The viability of our method is evaluated by considering a set of deep CNN architectures. We demonstrate that despite producing instrumentation drift during data collection between two modes of microscopy, this method performs comparably to similar segmentation tasks that used manual labeling. Additionally, we find that naïve pixel-wise segmentation results in small gaps and missing boundaries in the predicted grain boundary map. By incorporating topological information during model training, the connectivity of the grain boundary network and segmentation performance is improved. Finally, our approach is validated by accurate computation on downstream tasks of predicting the underlying grain morphology distributions which are the ultimate quantities of interest for microstructural characterization.

36 MATERIALS SCIENCE↗

Knowledge-matching based computational framework for genome-scale metabolic model refinement

Genome-scale metabolic models (GEMs) are mathematically structured knowledge base reconstructed from annotated genome of different organisms. With the advancement of next-generation sequencing technology, many organisms have had their genomes sequenced. However, obtaining a high-quality GEM is highly time-consuming, even with the introduction of several genome-scale reconstruction tools that offer automated draft network generation and gap filling. It has been recognized that the iterative process of manual curation and refinement is the limiting step of GEM development, and how to expedite the GEM refinement is still an open question. As cellular metabolism is a complex system with very high degree of freedom and redundancy, the principles and techniques developed in process systems engineering can be adapted to expedite GEM refinement. In this paper we present a knowledge-matching based computation framework for GEM refinement, and demonstrate the effectiveness of the proposed solution using the refinement of a GEM for Clostridium tyrobutyricum.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Tag you're it: Application of stable isotope labeling and LC-MS to identify the precursors of specialized metabolites in plants

Untargeted liquid chromatography/mass spectrometry (LC-MS) can contribute a comprehensive and unbiased picture of the metabolic space of plants. These data can be used to quantify natural metabolite variation for genome wide association studies, to compare global metabolic responses from environmental or genetic perturbations, and to identify previously undescribed metabolites in Nature. A major limitation with untargeted metabolomics is the classification and identification of the thousands of metabolite features that can be detected in a single analytical run. Isotopic labeling improves the informational value of these datasets by categorizing metabolites as being derived from specific upstream precursors and/or to known metabolic pathways. When a 13 C-labeled precursor is fed to either a plant or tissue, the downstream metabolites produced from it have a higher m/z value than the molecules in the pre-existing pool, generating an m/z peak pair that can be specifically identified within the MS data. In this paper, we outline methods and principles to consider when supplementing untargeted MS data with isotopic labeling, including how to choose the appropriate isotopic label, grow and feed plant tissues to maximize label uptake and incorporation into derivatives, optimize LC-MS methods, and interpret the resulting labeling data. Although the focus here is on annotation of amino acid-derived metabolites using LC-MS, we anticipate that the methods are generally adaptable to other precursors, plant species, and chromatographic approaches.

59 BASIC BIOLOGICAL SCIENCES↗

Scalable in situ non-destructive evaluation of additively manufactured components using process monitoring, sensor fusion, and machine learning

Laser Powder Bed Fusion (L-PBF) Additive Manufacturing (AM) is among the metal 3D printing technologies most broadly adopted by the manufacturing industry. However, the current industry qualification paradigm for critical-application L-PBF parts relies heavily on expensive non-destructive inspection techniques, which significantly limits the use-cases of L-PBF. In situ monitoring of the process promises a less expensive alternative to ex situ testing, but existing sensor technologies and data analysis techniques struggle to detect sub-surface flaws (e.g., porosity and cracking) on production-scale L-PBF printers. In this work, an in situ NDE (INDE) system was engineered to detect subsurface flaws detected in X-Ray Computed Tomography (XCT) directly from process monitoring data. A multilayer, multimodal data input allowed the INDE system to detect numerous subsurface flaws in the size range of 200–1000µm using a novel human-in-the-loop annotation procedure. Furthermore, a framework was established for generating probability-of-detection (POD) and probability-of-false-alarm (PFA) curves compliant with NDE standards by systematically comparing instances of detected subsurface flaws to post-build XCT data. Here, we also introduce for the first time in the AM in situ sensing literature the a 90/95 – the flaw size corresponding to a 90% detection rate on the lower 95% confidence interval of the POD curve. The INDE system successfully demonstrated POD capabilities commensurate with traditional NDE methods. Traditional ML performance metrics were also shown to be inadequate for assessing the ability of the INDE system’s flaw detection performance. It is the hope of the authors that future studies will adopt the POD and PFA approach outlined here to provide better insight into the utility of process monitoring for AM.

36 MATERIALS SCIENCE↗

Application of unsupervised deep learning to image segmentation and in-situ contact angle measurements in a CO 2 -water-rock system

Rock surface wettability is a critical property that regulates multiphase flows in porous media, which can be quantified using the surface contact angle (CA). X-ray micro-computed tomography (μCT) provides an effective approach to in-situ measurements of surface CAs. However, the CA measurement accuracy depends significantly on the quality of CT image segmentation, which is the clustering of CT pixels into separate phases. Inspired by this, we developed a deep learning (DL)-based CA measurement workflow. Motivated by the recent tremendous progress in unsupervised learning techniques and aiming to avoid expensive manual data annotations, an unsupervised DL pipeline for CT image segmentation was proposed and implemented, which includes unsupervised model training and post-processing. The unsupervised model training was driven by a novel loss function constrained with feature similarity and spatial continuity and implemented by iterative forward and backward paths; the former clustered the pixel-wise feature vectors extracted by convolution neural networks, whereas the latter updated the parameters using gradient descent. An over-segmentation strategy was adopted for model training. The post-processing steps based on agglomerative hierarchical clustering (AHC) were implemented to further merge the over-segmented model output to the desired cluster number, which is intended to improve the efficiency of image segmentation. The developed unsupervised DL pipeline was compared with other commonly-used image segmentation methods using pixel-wise and physics-based evaluation metrics on a synthetic raw-image dataset, which had a known ground truth. The unsupervised DL pipeline showed the best performance. Next, the segmented images were input to an automatic CA measurement tool, and the results were validated by comparisons with manual measurements. The CA values from the manual and automatic measurements showed similar distributions and statistical properties. The automatic measurement demonstrated a wider spectrum because of the much larger number of measurement data points. The primary novelty of the unsupervised DL pipeline developed in this study lies in the novel loss function and the over-segmentation strategy associated with AHC post-processing. Finally, the workflow has been proven an efficient tool for pore-scale wettability characterization, which has a wide range of applications in fundamental studies of multiphase flows in natural porous media, which have critical implications to geological carbon sequestration, hydrocarbon energy recovery, and contaminant transport in groundwater.

42 ENGINEERING↗

Capsule network-based semantic segmentation model for thermal anomaly identification on building envelopes

Thermography technology is widely used to inspect thermal anomalies in building façade systems. Computer vision-based techniques provide opportunities to autonomously detect such heat anomalies to significantly improve the efficiency of decision-making for building envelope retrofitting and maintenance. Here, in this work, we propose a novel Capsule Network-based deep learning model – CapsLab – that detects and identifies thermal anomalies by semantic segmentation. CapsLab is built based on our proposed prediction-tuning capsule (PT-Capsule) layer. Different from a traditional capsule layer, which consists of part-whole transformation and capsule-routing process, the proposed layer is composed of a prediction and tuning process, which helps decreasing the number of model parameters significantly. While the applicability of traditional Capsule Networks (CapsNets) has been limited to simpler tasks and smaller datasets due to their scalability issue, we can leverage the lightweight of the proposed PT-Capsule layer, and apply it to the semantic segmentation task. In this work, we also employ our previously presented performance metric, referred to as the Anomaly Identification Metric (AIM) (Kakillioglua et al. 2021), to evaluate the segmentation outputs. Traditional performance metrics do not accurately reflect the true performance of the segmentation models in thermal anomaly identification due to the high subjectivity in the annotation process and higher overlap ratio sensitivity of the standard metrics. AIM, on the other hand, is robust to these drawbacks. Experimental results show, both qualitatively and quantitatively, that our proposed segmentation method can effectively segment the thermal anomalies. Specifically, our model provides 9.38% and 13.53% improvements over the baseline model – DeepLabV3+ – based on traditional mIoU score and the AIM score, respectively, while requiring less model parameters and less computation at the same time. In addition, the scores that the AIM metric generates better align with the scores provided by building performance experts.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Multi-omics profiling of the cold tolerant Monoraphidium minutum 26B-AM in response to abiotic stress

Microalgae that are of interest for biofuel production must be able to tolerate environmental changes that occur in outdoor cultivation systems. While algal cultures may experience daily temperature fluctuations and seasonal environmental changes, the underlying mechanisms that control and regulate physiological responses and adaptation to environmental pressures are largely unknown. Systems-level characterization enabled by functional genomics can help identify biochemical pathways that promote stability and productivity of algae in various environmental conditions. Monoraphidium minutum 26B-AM, a freshwater green microalga, was identified as a top performer in biomass production in winter season screens. We sequenced the genome of M. minutum 26B-AM and applied our multi-omics pipeline to profile this high potential strain under high salt and cold temperature perturbations. Through comparative analysis, including other green algae in the class Chlorophyceae, we identified gene families unique to the genus Monoraphidium, including a desaturase that has been linked to cold tolerance in plants. We observed that osmolytes, such as trehalose, proline and betaine, accumulate under salt stress, coinciding with upregulation of genes involved in biosynthesis of these metabolites. From the genome annotation, we reconstructed a metabolic model to provide a detailed map of the metabolic pathways and can be used to simulate growth and reaction fluxes. This multi-omics analysis provides a foundation to explore algal strain potential for biofuel applications, guides strain engineering, and expands our understanding of metabolic and regulatory mechanisms of algae in applied systems.

59 BASIC BIOLOGICAL SCIENCES↗

What you get is not always what you see—pitfalls in solar array assessment using overhead imagery

Effective integration planning for small, distributed solar photovoltaic (PV) arrays into electric power grids requires access to high quality data: the location and power capacity of individual solar PV arrays. Unfortunately, national databases of small-scale solar PV do not exist; those that do are limited in their spatial resolution, typically aggregated up to state or national levels. While several promising approaches for solar PV detection have been published, strategies for evaluating the performance of these models are often highly heterogeneous from study to study. The resulting comparison of these methods for practical applications for energy assessments becomes challenging and may imply that the reported performance evaluations overly optimistic. The heterogeneity comes in many forms, each of which we explore in this work: the degree of diversity of the locations and sensors (e.g. different satellites, aerial photography) from which the training and validation data originate, the validation of ground truth (manual annotation of imagery vs known solar PV locations), the level of spatial aggregation (e.g. array-level vs regional estimates), and inconsistencies in the training and validation datasets (e.g. different datasets are used for each study and those data are not always made accessible). For each, we discuss emerging practices from the literature to address them or suggest directions of future research. As part of our investigation, we evaluate solar PV identification performance in two large regions: the entire state of Connecticut and the city of San Diego, CA. In Connecticut, we also use 33,114 known parcel-level solar PV installations from Berkeley Lab’s Tracking the Sun dataset to evaluate parcel-level performance and evaluate capacity estimates using 169 municipalities. We also make our code (which we call SolarMapper), pre-trained models, training data, and predictions publicly available and provide a web portal for interactively inspecting each prediction that was made. Here our findings suggest that traditional performance evaluation of the automated identification of solar PV from satellite imagery may be optimistic due to common limitations in the validation process. The takeaways from this work are intended to inform and catalyze the large-scale practical application of automated solar PV assessment techniques by energy researchers and professionals.

14 SOLAR ENERGY↗

Allosteric prediction via convolutional neural networks and protein structural and dynamical features

Allostery is the phenomenon whereby a binding event or covalent modification at one site in a protein modulates function at a distal site, thus changing a protein’s functional state. As such, it is a ubiquitous aspect of protein functional regulation. Computationally predicting allosteric states is important as part of the broader challenge of functional annotation, but it also has practical implications for drug development, as targeting an allosteric site often affords greater specificity compared with targeting an orthosteric site. This study introduces a machine learning approach to predict the allosteric functional state using the small G-protein KRas as the model system, due to its implication in many types of cancer and being well studied as a result with many x-ray crystallographic structures of KRas available with different mutations and ligands bound. Using structural and dynamical features that can be cast as images, namely interatomic distances, contact maps, covariance, and mutual information, supervised learning was performed using convolutional neural networks. Two pretrained convolutional neural network architectures, GoogLeNet and ResNet18, were fine-tuned to classify KRas into active or inactive states based on these features. Across training regimes, atomic contact maps emerged as the most effective structural feature, whereas linearized mutual information outperformed covariance in capturing dynamical correlations relevant to allostery. Models achieved significant validation accuracy, with atomic contact maps yielding up to 90% accuracy. In conclusion, the findings suggest that integrating global structural rearrangements and correlated motion patterns with deep learning can reliably predict protein allosteric states, offering a promising framework for understanding allosteric regulation and developing targeted therapeutics.

Rajeshwar T., Rajitha [Oak Ridge National Laborato↗