Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Python codes”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Evidence-based Graph Adversary Mapping (EGRAM) [Poster]

Cybersecurity companies such as CrowdStrike, Dragos, Microsoft and Unit 42 categorize Advanced Persistent Threats (APTs) using their own naming schemes. As a result, these APTs are mapped to different malware sources and campaigns, all from differing sources, leading to inconsistent mapping. Inconsistent mapping causes confusion and adds further obscurity around these groups, making it difficult to track and mitigate APT cyberattacks. The Evidence-based Graph Adversary Mapping (EGRAM) tool remediates the mapping challenge by collecting, updating and converting adversary data and their sources into a valid, codified STIX v2.1 bundle which is then stored in a Neo4j graph database. It utilizes graph traversal methods and centrality analysis to generate actionable information as a Structured Threat Intelligence Graph (STIG), based on user queries. EGRAM exists as Python code and a Jupyter Notebook that acts as a searchable, evidence-based, source of intelligence for APT groups’ artifacts and cyber campaigns.

24 - POWER TRANSMISSION AND DISTRIBUTION↗

ORT: a workflow linking genome-scale metabolic models with reactive transport codes

Abstract Motivation Nutrient and contaminant behavior in the subsurface are governed by multiple coupled hydrobiogeochemical processes which occur across different temporal and spatial scales. Accurate description of macroscopic system behavior requires accounting for the effects of microscopic and especially microbial processes. Microbial processes mediate precipitation and dissolution and change aqueous geochemistry, all of which impacts macroscopic system behavior. As ‘omics data describing microbial processes is increasingly affordable and available, novel methods for using this data quickly and effectively for improved ecosystem models are needed. Results We propose a workflow (‘Omics to Reactive Transport—ORT) for utilizing metagenomic and environmental data to describe the effect of microbiological processes in macroscopic reactive transport models. This workflow utilizes and couples two open-source software packages: KBase (a software platform for systems biology) and PFLOTRAN (a reactive transport modeling code). We describe the architecture of ORT and demonstrate an implementation using metagenomic and geochemical data from a river system. Our demonstration uses microbiological drivers of nitrification and denitrification to predict nitrogen cycling patterns which agree with those provided with generalized stoichiometries. While our example uses data from a single measurement, our workflow can be applied to spatiotemporal metagenomic datasets to allow for iterative coupling between KBase and PFLOTRAN. Availability and implementation Interactive models available at https://pflotranmodeling.paf.subsurfaceinsights.com/pflotran-simple-model/. Microbiological data available at NCBI via BioProject ID PRJNA576070. ORT Python code available at https://github.com/subsurfaceinsights/ort-kbase-to-pflotran. KBase narrative available at https://narrative.kbase.us/narrative/71260 or static narrative (no login required) at https://kbase.us/n/71260/258. Supplementary information Supplementary data are available at Bioinformatics online.

54 ENVIRONMENTAL SCIENCES↗

PyOMP: Multithreaded Parallel Programming in Python

We know that Python is a widely used language in scientific computing. When the goal is high performance, however, Python lags far behind low-level languages such as C and Fortran. To support applications that stress performance, Python needs to access the full capabilities of modern CPUs. That means support for parallel multithreading. In this paper, we describe PyOMP, a system that enables OpenMP in Python. Programmers write code in Python with OpenMP, Numba generates code that compiles to LLVM, and the resulting programs run with performance that approaches that from code written with C and OpenMP. In this paper we provide an update on the PyOMP project and explain how to install it and use it to write parallel multithreaded code in Python.

97 MATHEMATICS AND COMPUTING↗

Data on Cu- and Ni-Si-Mn-rich solute clustering in a neutron irradiated austenitic stainless steel

The data presented in this article is supplementary to the research article “Phase instabilities in austenitic steels during particle bombardment at high and low dose rates” (Levine et al.). Needle-shaped samples were prepared with focused ion beam milling from a 304L stainless steel that was irradiated with fast neutrons (E 0.1 MeV) in the BOR-60 reactor at 318 °C to 47.5 dpa. Atom probe tomography (APT) experiments in voltage mode were then conducted on a Cameca LEAP 5000X HR. Atom position, range, and mass spectrum files after reconstruction with Cameca’s IVAS software are included. Cu- and Ni-Si-Mn-rich solute nanoclusters were identified and analyzed using the Open Source Characterization of APT Reconstructions (OSCAR) program. Python code for OSCAR, information on the program’s underlying algorithm, and sample output files are provided. A proximity histogram of a Ni-Si-Mn-rich cluster and a 1D density/solute concentration profile of a Cu-rich cluster are given to demonstrate OSCAR’s analytical functionalities. The provided APT dataset is valuable for benchmarking phase instabilities in neutron-irradiated austenitic stainless steels that occur at high doses. The OSCAR program can be reused to process other APT data sets where solute nanoclustering is of interest.

42 ENGINEERING↗

Data and scripts from: “Denoising autoencoder for reconstructing sensor observation data and predicting evapotranspiration: noisy and missing values repair and uncertainty quantification”

This data package includes data and scripts from the manuscript “Denoising autoencoder for reconstructing sensor observation data and predicting evapotranspiration: noisy and missing values repair and uncertainty quantification”.The study addressed common challenges faced in environmental sensing and modeling, including uncertain input data, missing sensor observations, and high-dimensional datasets with interrelated but redundant variables. Point-scaled meteorological and soil sensor observations were perturbed with noises and missing values, and denoising autoencoder (DAE) neural networks were developed to reconstruct the perturbed data and further predict evapotranspiration. This study concluded that (1) the reconstruction quality of each variable depends on its cross-correlation and alignment to the underlying data structure, (2) uncertainties from the models were overall stronger than those from the data corruption, and (3) there was a tradeoff between reducing bias and reducing variance when evaluating the uncertainty of the machine learning models.This package includes:(1) Four ipython scripts (.ipynb): “DAE_train.ipynb” trains and evaluates DAE neural networks, “DAE_predict.ipynb” makes predictions from the trained DAE models, “ET_train.ipynb” trains and evaluates ET prediction neural networks, and “ET_predict.ipynb” makes predictions from trained ET models.(2) One python file (.py): “methods.py” includes all user-defined functions and python codes used in the ipython scripts.(3) A “sub_models” folder that includes five trained DAE neural networks (in pytorch format, .pt), which could be used to ingest input data before being fed to the downstream ET models in ‘ET_train.ipynb” or ‘ET_predict.ipynb’.(4) Two data files (.csv). Daily meteorological, vegetation, and soil data is in “df_data.csv”, where “df_meta.csv” contains the location and time information of “df_data.csv”. Each row (index) in “df_meta.csv” corresponds to each row in “df_data.csv”. These data files are formatted to follow the data structure requirements and be directly used in the ipython scripts, and they have been shuffled chronologically to train machine learning models. The meteorological and soil data was collected using point sensors between 2019-2023 at(4.a) Three shrub-dominated field sites in East River, Colorado (named “ph1”, “ph2” and “sg5” in “df_meta.csv”, where “ph1” and “ph2” were located at PumpHouse Hillslopes, and “sg5” was at Snodgrass Mountain meadow) and(4.b) One outdoor, mesoscale, and herbaceous-dominated experiment in Berkeley, California (named “tb” in “df_meta.csv”, short for Smartsoils Testbed at Lawrence Berkeley National Lab).- See "df_data_dd.csv" and "df_meta_dd.csv" for variable descriptions and the Methods section for additional data processing steps. See "flmd.csv" and "README.txt" for brief file descriptions.- All ipython scripts and python files are written in and require PYTHON language software.

54 ENVIRONMENTAL SCIENCES↗

Cardinal: Seismic and Geoacoustic Array Processing

Data collected via seismic and infrasound array deployments are leveraged in the geosciences to detect and characterize a myriad of natural and anthropogenic sources. These deployments consist of numerous sensors placed in a predetermined configuration to amplify signal strength and improve the efficacy of array processing techniques used to measure signal directionality and waveform coherence. High‐fidelity feature extraction is often predicated on interstation distance as well as the frequency content and wavelength of an incident signal. Numerous array processing softwares analyze data in sequential frequency bands to obtain a more detailed characterization of a signal. However, current algorithms are limited in their ability to determine optimal array configuration for each band. We introduce an open‐source Python code, called Cardinal, to process seismic and infrasound array data in discretized time–frequency space with the option of applying an adaptive array design to determine optimal subarray configuration for each frequency band. To reduce computational time, the array processing step can be run in parallel using multithreading. Furthermore, the software has the capability to aggregate array processing results from different time–frequency pixels to produce separate sets of detections, or families, with added utility via the application of an adaptive semblance threshold, which aids in isolating signals‐of‐interest from coherent background noise. Upon appropriate configuration, Cardinal exhibits the potential to combine distinct seismic and infrasound phases into separate families.

Adaptive Array↗

Combining Astrometry and Elemental Abundances: The Case of the Candidate Pre-Gaia Halo Moving Groups G03-37, G18-39, and G21-22

While most moving groups are young and nearby, a small number have been identified in the Galactic halo. Understanding the origin and evolution of these groups is an important piece of reconstructing the formation history of the halo. Here we report on our analysis of three putative halo moving groups: G03-37, G18-39, and G21-22. Based on Gaia EDR3 data, the stars associated with each group show some scatter in velocity (e.g., Toomre diagram) and integrals of motion (energy, angular momentum) spaces, counter to expectations of moving-group stars. We choose the best candidate of the three groups, G21-22, for follow-up chemical analysis based on high-resolution spectroscopy of six presumptive members. Using a new Python code that uses a Bayesian method to self-consistently propagate uncertainties from stellar atmosphere solutions in calculating individual abundances and spectral synthesis, we derive the abundances of α- (Mg, Si, Ca, Ti), Fe-peak (Cr, Sc, Mn, Fe, Ni), odd-Z (Na, Al, V), and neutron-capture (Ba, Eu) elements for each star. We find that the G21-22 stars are not chemically homogeneous. Based on the kinematic analysis for all three groups and the chemical analysis for G21-22, we conclude the three are not genuine moving groups. The case for G21-22 demonstrates the benefit of combining kinematic and chemical information in identifying conatal populations when either alone may be insufficient. Comparing the integrals of motion and velocities of the six G21-22 stars with those of known structures in the halo, we tentatively associate them with the Gaia-Enceladus accretion event.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Development of the Intelligent, Preventive Infrared (IR) Inspection System Housed in Hybrid Robotic Platforms

Robots and robotic systems that are designed for inspection, environmental study, and health and safety aid are becoming an increasing necessity. However, there is a number of challenges that accompany robots that are designed for these specific applications. These challenges include: navigating compact, enclosed spaces, travelling over multiple terrains and large obstacles, using the proper sensing and detection methods to assess an environment, and the use of lightweight and durable materials. The robotic platforms currently in development look at all of these challenges and attempt to overcome them. These designs specific use of hybrid robotic platforms, or platform that utilizes soft and rigid materials, allows for a more flexible platform and makes environments more navigable. To further improve the navigation of the platforms and environmental assessment, a novel infrared detection system is housed in the platforms to create a robot that can be used for the applications listed above and more. Objectives: Further develop two types of robotic platforms that utilize additive manufacturing, soft materials, and rigid materials. Continue the development of an intelligent and inhibitory infrared detection system based on an artificial intelligence (AI) algorithm. Continued study and fabrication of active soft materials designed for both sensing and actuation in hybrid robotic systems. Improve additive manufacturing fabrication to design rigid and semi-rigid components for hybrid robotic platforms. Transformable Wheel Robotic Platform: The chassis, wheels, tires and inspection system housing use different additive manufacturing techniques for fabrication. Continued work with additive manufacturing has lead to studies in metal-based printing and modular design and manufacturing. The new platform design with integrated electrical component printed. This will allow integration of the sensor housing onto the platform. Electrical components are being tested for battery life and performance. To improve this performance, such as integration of Lithium Polymer (LiPo) batteries. Snake Robotic Platform: The main focus of the development has centered around liquid-based soft actuators. That act on the principles of electrostatic and hydraulic actuation. A liquid dielectric sits between two compliant electrodes, contained by a flexible polymer shell. The electrodes and film gradually collapse toward each other from one corner of the electrode to the other. When the electrodes and film close together, a majority of the fluid is pushed into the area not covered by an electrode. A thin layer of the liquid dielectric remains between the electrode. The actuators will be stacked to cause large displacement, and move the linkages. The chassis of this platform uses purely additively manufactured linkages. Intelligent, Preventive IR Inspection System: Development of the AI for the system has lead to using a Scikit-Learn which assists in creating predictive models based on Regression, clustering, classification etc. To improve the infrared thermometry for low emissivity sources, work on the fabrication of a tandem photoconductive infrared thermometer was a main focus. Distance-Voltage-Temperature response data has been collected in the range of 7 cm - 100 cm and 200-400 deg. C. Modifications were made to the existing test bench to have a better control over the data. Automated data collection was realized using a Python code and an Arduino controlled stepper motor to increase the sample rate. A protective enclosure has been built around the setup to minimize the effect of the environment on the measurements. To better predict temperature, different AI models are being optimized. The regression model, LARS showed a high accuracy but had convergence issues and only works for the current test set-up. Results: The Transformable Wheel Robot has developed into a more flexible and modular platform. With the improvements to the current work, effort on the tire or soft gripper design been a large focus. The soft grippers will be interchange able to allow for increased performance in identified terrain types. The development of the liquid-based actuators allows the snake robotic platform to achieve the goals of being flexible and able to navigate confined spaces. However, there is room for improvement. Optimization work is currently being done in COMSOL Multiphysics. With the current IR system set-up the LARS model perfectly predicts the data; however, considering the mobility aspect of the project other models will allow for an optimized system. Testing of other model types is currently being done.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

Computational and Experimental Investigation of Chiral and Achiral Two‐Dimensional Organic Lead Bromide Perovskites: Octahedral Distortions and Electronic and Optical Properties

A computational investigation is presented, in conjunction with synthesis and experimental characterization, into the structural, electronic, and optical properties of layered two-dimensional organic lead bromide perovskites. Materials based on the chiral (R/S)-4-fluoro-α-methylbenzylammonium (R/S-FMBA), which have been shown to lead to bright room-temperature circularly polarized luminescence, are contrasted with the similar achiral 4-fluorobenzylammonium (FBA). Using density functional theory (DFT) with van der Waals (vdW) corrections, relaxed structures (compared with X-ray diffraction, XRD) and optical absorption spectra (compared with experiments) are studied, as well as band structure and orbital character of transitions. A Python code is developed and provided to calculate octahedral distortions and compare DFT and XRD results, finding that vdW corrections are important for accuracy and that DFT overestimates octahedral tilt angles. (FMBA) 2 PbBr 4 shows among the largest tilt angle differences (often termed Δ β ) reported, 14°–15°, indicating strong inversion symmetry-breaking, which enables its chiral emission. A large resulting Dresselhaus spin-splitting effect is found. The lowest-energy optical transitions involve the perovskite only and are polarized within the layer. This work furthers understanding of structure-property relations with applications to optoelectronics and spintronics.

UV/vis spectroscopy↗

Gap-filling eddy covariance methane fluxes: Comparison of machine learning model predictions and uncertainties at FLUXNET-CH4 wetlands

Time series of methane fluxes measured by eddy-covariance require gap-filling to estimate annual emissions. Gap-filling methane fluxes is challenging because of high variability and complex responses to multiple drivers. To date, there is no widely established gap-filling standard for methane, with regards both to the best model algorithms and predictors. In this study, we address the need for standardization by synthesizing results of gap-filling methods applied at 17 wetland sites spanning boreal to tropical regions including all major wetlands classes and two rice paddies. We introduce new procedures for: 1) creating realistic artificial gap scenarios, 2) training and evaluating gap-filling models without overstating performance, and 3) predicting half-hourly methane fluxes and annual emissions with robust uncertainty estimates. We tested a conventional method (marginal distribution sampling) and four machine learning algorithms - penalized linear regression, artificial neural networks, random forests, and boosted decision trees - and four predictor sets, including temporal, meteorological, ecosystem carbon and energy flux, and soil predictors. We find that the conventional method can achieve similar median performance to the machine learning models but is worse than the best machine learning models and relatively insensitive to predictor choices. Of the machine learning models, decision tree algorithms performed the best in cross-validation experiments, even with a baseline predictor set, and artificial neural networks showed comparable performance when using all predictors. Soil temperature was frequently the most important predictor whilst water table depth was important at sites with substantial water table fluctuations, highlighting the value of data on soil conditions. Raw gap-filling uncertainties from the machine learning models were underestimated and we propose a method to calibrate uncertainties to observations. Finally, we gap-fill and provide summary evaluation metrics for all 81 sites in the FLUXNET-CH4 community dataset and publicly release the python code for model development, evaluation, and uncertainty estimation.

42 ENGINEERING↗

Solution and sensitivity analysis of nonlinear equations using a hypercomplex-variable Newton-Raphson method

Here, the classical Newton-Raphson (NR) method for solving nonlinear equations is enhanced in two ways through the use of hypercomplex variables and algebra. In particular, i) the Jacobian is computed in a highly accurate and automated way, and ii) the derivative of the solution to the nonlinear equations is computed with respect to any parameter contained within the system of equations. These advances provide two significant enhancements in that it is straightforward to provide an accurate Jacobian and to construct a reduced order model (ROM) of arbitrary order with respect to any parameter of the system. The ROM can then be used to approximate the solution for other parameter values without requiring additional solutions of the nonlinear equations. Several case studies are presented including 1D and 2D academic examples with fully functioning Python code provided. Additionally, a case of study of the catenary of an elastic cable subject to its own weight and a vertical point load. Derivatives up to 10th order were computed with respect to material, loading, and geometrical parameters. The derivatives were used to generate reduced order models of the cable deformation and reaction forces at its ends with respect to multiple input parameters. Results show that from a single hypercomplex evaluation of the cable under a single vertical point load, it is possible to generate an accurate reduced order model capable of predicting the cable deformation with 1.5 times the load in the opposite direction and with 3.5 times the load in the same direction without resolving the system of equations.

97 MATHEMATICS AND COMPUTING↗

Graph neural networks for CO 2 solubility predictions in Deep Eutectic Solvents

Deep Eutectic Solvents (DESs) are a promising class of solvents for CO 2 capture. DESs are complex mixtures that can be designed to optimize CO solubility and overall capture process efficiency. However, the vast design landscape of DES mixtures makes experimental investigation prohibitive; as such, there is a need for computational models that can quickly and efficiently navigate the design space and inform data collection efforts. In this work, we propose Graph Neural Network (GNN) models for predicting CO 2 solubility for DESs; the GNN leverages a mixture graph representation that captures the molecular structure of the DES components as well as their intermolecular interactions. Here, we compare the GNN framework against alternative architectures (neural networks, graph convolution networks, and random forests) and data representations (molecular fingerprints, sigma profiles, and graphs). We show that the proposed approach offers superior predictive performance; specifically, we show that solubility can be predicted reliably directly from molecular structure (without the need of using sigma profiles as proposed in previous studies). This result is important, as obtaining sigma profiles requires expensive density functional theory computations. We also explored the ability of GNNs to predict solubility for new DES mixtures and operating conditions. We found that the model extrapolates across temperature reliably. However, we also found deficiencies in the ability of the model to predict solubility for DES mixtures, pressures, and molar ratio not included in the training sets; we show that this is due to an inherent lack of chemical diversity in datasets available in the literature. The proposed computational capabilities can thus help navigate the design space of DES and inform data collection efforts. Our models, data, and benchmarks are shared as Python code implemented in Jupyter notebooks.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Short-term electricity load forecasting: Application-driven evaluation of machine learning models across spatial and temporal scales

As we transition towards a decarbonized economy, the integration of variable renewable energy resources and new demands (e.g., electric vehicles, heat pumps) into the electricity grid places unprecedented pressure on grid operators to effectively anticipate and manage peak load. In this context, machine learning algorithms are proving to be indispensable for accurate short-term load forecasting, a crucial task to address these challenges. This study benchmarks 6 machine learning algorithms, including three neural networks and three tree-based algorithms, across various levels of spatial aggregation and time horizons (1, 4, 8, 24, and 48 h). The central contribution of this work is the comparison and analysis of load forecasting models not only based on statistical metrics, but also based on a novel error metric, which evaluates the cost implications of forecast errors for power system stakeholders. Results show that tree-based models outperform neural networks, based on statistical metrics, and yield less skewed error distributions for most spatial scales. However, through the lens of the novel error metric, neural networks are the more competitive choice, especially for forecast horizons that exceed 8 h. The study concludes with actionable recommendations to grid operators and highlights the need for the development of error metrics that link forecasting accuracy to operational costs. To promote transparency and open science, the datasets and Python code are open-sourced via a supplementary repository.

Houben, Nikolaus↗

Economic assessment of seismic monitoring for underground hydrogen storage

Underground hydrogen storage (UHS) plays a key role in the energy landscape. However, like other subsurface engineering technologies, UHS may cause leakage into the groundwater or atmosphere and possibly induce local seismicity. To reduce these risks, seismic monitoring could be a viable technique to track the UHS plume, detect leakages, and locate induced seismicity events. Seismic monitoring has been proposed to safely monitor UHS, but research in this area is still new and requires field studies. Lab and theoretical studies have demonstrated the validity of seismic monitoring for UHS. Therefore, it is imperative to analyze the economic feasibility of seismic monitoring for UHS. Hence, we develop a cost model and open-source Python code for seismic monitoring that considers types of seismometers, comprehensive operational scenarios, detection thresholds, and long-term leakage monitoring. A case study is further provided to validate the cost model on reservoir simulations of UHS. We find that the levelized cost for a 10-year operating UHS site will range on the order of ∼0.003 $\$$/kg. The methods developed in this study could also be applied to the monitoring of groundwater, gas, and/or wastewater injection.

08 HYDROGEN↗

GMFOLD: Subgraph matching for high-throughput DNA-aptamer secondary structure classification and machine learning interpretability

Aptamers are oligonucleotide receptors that bind to their targets with high affinity. Here, we consider aptamers comprised of single-stranded DNA that undergo target-binding-induced conformational changes, giving rise to unique secondary and tertiary structures. Given a specific aptamer primary sequence, there are well-established computational tools (notably mfold) to predict the secondary structure via free energy minimization algorithms. While mfold generates secondary structures for individual sequences, there is a need for a high-throughput process whereby thousands of DNA structures can be predicted in real-time for use in an interactive setting, when combined with aptamer selections that generate candidate pools that are too large to be experimentally interrogated. We developed a new Python code for high-throughput aptamer secondary structure determination (GMfold). GMfold uses subgraph matching methods to group aptamer candidates by secondary structure similarities. We also improve an open-source code, SeqFold, to incorporate subgraph matching concepts. We represent each secondary structure as a lowest-energy bipartite subgraph matching of the DNA graph to itself. These new tools enable thousands of DNA sequences to be compared based on their secondary structures, using machine-learning algorithms. This process is advantageous when analyzing sequences that arise from aptamer selections via systematic evolution of ligands by exponential enrichment (SELEX). This work is a building block for future machine-learning-informed DNA-aptamer selection processes to identify aptamers with improved target affinity and selectivity and advance aptamer biosensors and therapeutics.

Aptamer↗

Evaluating wind speed and power forecasts for wind energy applications using an open-source and systematic validation framework

Building on the verification and validation work developed under the Second Wind Forecast Improvement Project, this work exhibits the value of a consistent procedure to evaluate wind power forecasts. We established an open-source Python code base tailored for wind speed and wind power forecast validation, WE-Validate. The code base can evaluate model forecasts with observations in a coherent manner. To demonstrate the systematic validation framework of WE-Validate, we designed and hosted a forecast evaluation benchmark exercise. We invited forecast providers in industry and academia to participate and submit forecasts for two case studies. We then evaluated the submissions with WE-Validate. Our findings suggest that ensemble means have reasonable skills in time series forecasting, whereas they are often inferior to single ensemble members in wind ramp forecasting. Adopting a voting scheme in ramp forecasting that allows ensemble members to detect ramps independently leads to satisfactory skill scores. Throughout this document, we also emphasize the importance of using statistically robust and resistant metrics as well as equitable skill scores in forecast evaluation.

17 WIND ENERGY↗

Identification of Serine-Containing Microcystins by UHPLC-MS/MS Using Thiol and Sulfoxide Derivatizations and Detection of Novel Neutral Losses

Microcystins (MCs) are hepatotoxic cyclic heptapeptides produced by cyanobacteria, and their structural diversity has led to the discovery of more than 300 congeners to date. However, with known amino acid combinations, many more MC congeners are theoretically possible, suggesting many remain unidentified. Herein, two novel serine (Ser)-containing MCs were putatively identified in a Lake Erie cyanobacterial harmful algal bloom (cyanoHAB), using high-resolution UHPLC-MS as well as thiol and sulfoxide derivatization procedures. These MCs contain an α,β-unsaturated carbonyl on methyl dehydroalanine (Mdha) residue that undergoes Michael addition to produce a thiol-derivatized MC. Derivatization reactions using various thiolation reagents were followed by MS/MS, and two Python codes were used for data analysis and structural elucidation of MCs. Two novel MCs containing Ser at position 1 (i.e., next to Mdha) were putatively identified as [Ser 1 ]MC-RR and [Ser 1 ]MC-YR. Using thiol- and sulfoxide-modified [Ser 1 ]MCs, identifications were confirmed by the observation of specific neutral losses of the oxidized thiols or sulfoxides in CID-MS/MS spectra in both positive and negative electrospray ionization (ESI) modes. These novel neutral losses are unique for MCs with Mdha and an adjacent Ser residue. In conclusion, data suggest that a gas-phase reaction occurs between oxygen from adjacent Ser residue and sulfur of the Mdha-bonded thiol or sulfoxide, which leads to the formation and detection of stable cyclic MC ions in MS/MS spectra at m/z values corresponding to the loss of oxidized thiols or oxidized sulfoxides from Ser 1 -containing MCs.

Premathilaka, Sanduni H.↗

Machine Learning for Materials Scientists: An Introductory Guide toward Best Practices

This Methods/Protocols article is intended for materials scientists interested in performing machine learning-centered research. Herein, we cover broad guidelines and best practices regarding the obtaining and treatment of data, feature engineering, model training, validation, evaluation and comparison, popular repositories for materials data and benchmarking data sets, model and architecture sharing, and finally publication. In addition, we include interactive Jupyter notebooks with example Python code to demonstrate some of the concepts, workflows, and best practices discussed. Overall, the data-driven methods and machine learning workflows and considerations are presented in a simple way, allowing interested readers to more intelligently guide their machine learning research using the suggested references, best practices, and their own materials domain expertise.

36 MATERIALS SCIENCE↗