Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data analysis methods”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

Machine learning to discover mineral trapping signatures due to CO 2 injection

Mineral trapping is pursued as a geological CO 2 sequestration (GCS) mechanism because it permanently stores CO 2 in solid phases or minerals. However, CO 2 mineral-trapping mechanisms are poorly understood due to (1) lack of sufficient field and laboratory data characterizing these complex processes, and (2) challenges to develop site-specific reactive-transport models coupling fluid flow and geochemical reactions occurring at various temporal (from milliseconds to years) and spatial (from pore (millimeters) to field (kilometers)) scales. Reactive transport with additional complexities such as heterogeneity can make the simulation outputs even more difficult to interpret because of complex nonlinearity and multi-scale interdependencies. Furthermore, the values of model outputs such as concentrations can vary by several orders of magnitude, making it harder to correlate and characterize the impact of the variables via traditional data interpretation techniques such as exploratory data analyses. Recently, machine learning (ML) has shown promise in feature discovery and in highlighting hidden mechanisms that cannot be obtained by existing data-analytics and statistical methods. In this study, we applied an unsupervised ML approach, non-negative matrix factorization with custom -means clustering (NMF) to the data generated by reactive-transport simulations of GCS. The reactive-transport data consisted of 19 attributes, including four physio-chemical variables (pH, porosity, aqueous CO 2 , and sequestered CO 2 ), six chemical species (K + , Na + , HCO, Ca 2+ , Mg 2+ , Fe 2+ ), and four carbonate minerals (calcite, dolomite, siderite, and ankerite), a feldspar mineral (albite), and four clay minerals (illite, clinochlore, kaolinite, and smectite) over a period of 200 years of simulation time. Furthermore, the simulation data used was for Morrow B sandstone at the Farnsworth hydrocarbon unit in Texas. Data are sampled at two locations within the model domain: (1) at the injection well and (2) 200 m west of the injection well. The injection was performed for a period of 10 years. Using NMF, we estimated the temporal interdependencies among the 19 attributes over a span of 200 years. We found that NMF was able to identify four reaction stages and their dominant attributes; these cannot be directly discerned through traditional visualization (e.g., line plots, Pareto analysis, Glyph-based visualization methods) or exploratory data analysis tools of the simulation data. The four stages were: reactions in the injection phase followed by short-, mid-, and long-term reactions. The NMF analysis also revealed that 10 among the 19 attributes are dominant. These dominant attributes for mineral trapping include calcite, dolomite at injection well, siderite at 200 m away from the injection well, clinochlore, kaolinite, Na + , K + , Ca 2+ , Mg 2+ , pH, and aqeuous CO 2 . Finally, at late times (65–200 years), our results showed that calcite plays a major role in mineral trapping with insignificant contribution from siderite, ankerite, and clay minerals. These findings make the proposed unsupervised ML-model attractive for reactive-transport sensing towards real-time GCS monitoring.

54 ENVIRONMENTAL SCIENCES↗

Global analysis peak fitting for chemical spectroscopy data

The present invention relates to methods for analyzing a chemical sample. For instance, the methods herein allow for global analysis of spectroscopy data in order to extract useful chemical properties from complicated multidimensional data. Such analysis can optionally employ data compression to further expedite computer-implemented computation. In particular, the methods herein provide global analysis of data matrices explained by both linear and non-linear terms.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Practical Guide to Chemometric Analysis of Optical Spectroscopic Data

The methodology and mathematical treatment of several classic multivariate methods for the analysis of spectroscopic data is demonstrated in a straightforward way that can be used as a basis for teaching an undergraduate introductory course on chemometric analysis. The multivariate techniques of classical least squares (CLS), principal component regression (PCR), and partial least squares (PLS), as well as the univariate Beer’s law method have been described and compared, building students’ understanding by starting with the univariate method and progressing step by step into the multivariate methods. Equations for the production of regression vectors from training set spectral data is described and their use demonstrated for the prediction of constituent concentrations on a separate validation set of spectra. Extreme care is taken to ensure consistency in variable formatting of data matrices. This provides a key foundation to understanding how spectral data are manipulated using these different mathematical approaches for building quantitative regression models. Each method is applied to a real-world data set, and the results are discussed to show students the types of information that can be gleaned from each method. A training set comprised of 20 infrared absorbance spectra containing 3 constituents (benzene, polystyrene, and gasoline) of known composition are used to demonstrate the matrix operations for each regression method. A separate set of 12 real-world napalm samples (containing benzene, polystyrene and gasoline) are used as a validation set to demonstrate the ability to utilize the regression models on an unknown dataset. A toolbox (PNNL Chemometric Toolbox) written in MATLAB language is supplied in the Supplemental Information file and can be used as a companion for understanding the development and deployment of the chemometric algorithms described in this paper. The datasets of the infrared spectra are also supplied, allowing users to build and inspect the chemometric models on their own. Finally, the Toolbox includes scripts to assist users in loading their own datasets into MATLAB and performing CLS, PCR, and PLS on their data.

Upper-Division Undergraduate, Analytical Chemistry↗

Niche-DE: niche-differential gene expression analysis in spatial transcriptomics data identifies context-dependent cell-cell interactions

Existing methods for analysis of spatial transcriptomic data focus on delineating the global gene expression variations of cell types across the tissue, rather than local gene expression changes driven by cell-cell interactions. We propose a new statistical procedure called niche-differential expression (niche-DE) analysis that identifies cell-type-specific niche-associated genes, which are differentially expressed within a specific cell type in the context of specific spatial niches. We further develop niche-LR, a method to reveal ligand-receptor signaling mechanisms that underlie niche-differential gene expression patterns. Niche-DE and niche-LR are applicable to low-resolution spot-based spatial transcriptomics data and data that is single-cell or subcellular in resolution.

59 BASIC BIOLOGICAL SCIENCES↗

Development of Analysis Methods that Integrate Numeric and Textual Equipment Reliability Data

Within the Light Water Reactor Sustainability (LWRS) program, the Risk-Informed Systems Analysis (RISA) Pathway is performing collaborative research on the development and deployment of technologies designed to assist operating nuclear power plants (NPPs) to reduce operating costs improve plant reliability and availability. One of the RISA research areas is focusing on the development of methods and tools designed to optimize plant operations (e.g., maintenance/replacement schedules, optimal maintenance postures for plant structures, systems, and components [SSCs]) in a manner that is more cost effective than current approaches and makes better use of available SSC health data. The Risk-Informed Asset Management (RIAM) project targets this research area by creating a direct bridge between component equipment reliability (ER) data and system engineer decision making regarding maintenance activity scheduling and component aging management. In this respect, one challenge that NPP system engineers are facing is that the amount of ER data being continuously generated is not only extremely large in size, but it comes in different forms: textual (e.g., condition or maintenance reports) and numeric (e.g., generated by monitoring systems). All these data elements provide them with valuable insights and information regarding: 1) the discovery of anomalous behaviors or degradation trends, 2) the identification of the possible causes behind such behaviors/trends, and 3) the prediction of their direct consequences. However, several challenges have proved to be roadblocks to this process. While some of these challenges are technical in nature (i.e., data are often distributed over several physical servers/databases), others are conceptual in nature: data elements come in different formats (e.g., numeric or textual), and measured values have different scales (e.g., vibration spectra and oil temperature). The activities performed by the RIAM project during FY23 directly tackles the need to simultaneously integrate the analysis of ER data in all its forms, numeric and textual. Note that such task has never been performed before due to the complexity of the systems under consideration but, most importantly, because of the technical challenges behind the harmonization of ER data formats and the lack of adequate computational methods to analyze them. Our approach borrows ideas and concepts from the medical field where integration of several data sources is vital to assist medical practitioners to perform correct diagnosis and indicate optimal treatments. In our view a NPP asset is equivalent to a patient in a medical context. The main difference is the complexity of a human body is a magnitude more complex when compared to typical assets commonly present in NPPs (e.g., centrifugal pumps, or motor operated valves). This simplifies our first requirement when analyzing heterogenous ER data formats: to put data into “context”. Context is here intended as the additional piece of information that is needed by ER data analysis tools to understand what these data elements are referring to, i.e., which king of knowledge they are generating. In our context, this knowledge can be translated into models that capture the form and functional architecture of assets/systems, their dependencies, and how they interact. These models actually emulate the knowledge that that NPP system engineers possess about assets and systems; this is their key of success when analyzing ER data, their challenge is ability to handle large amount of data. Here, we employ model-based system engineering (MBSE) models of systems and assets to represent and capture their architecture and functional, i.e. cause-effect, relations. Then, ER data elements are processed by identifying first of all which elements of the developed MBSE elements they are referring to. For numeric ER data this task is fairly easy since it is possible to precisely pinpoint what MBSE elements the corresponding sensor are observing (e.g., bearing temperature of a centrifugal pump). Task is much harder for textual data since the information contained in issue or maintenance reports needs to “be understood” by a computational tool. Here we called this process as “knowledge extraction”. Once again, we borrow the experience in the medical field where methods to extract knowledge from textual data have been developed in the past decade. The missing element for us is the availability of a complete dictionary of NPP related concepts (in addition to the MBSE models presented earlier) that can put “text into context”. In FY23, such dictionary has been developed along with all the computational elements required for knowledge extraction. Lastly, once numeric and textual ER data elements have been processed and “understood”, then the last step is the discovery of possible cause-effect relations among them. This is performed by observing if a logical connection through the MBSE models exists, and if the

97 MATHEMATICS AND COMPUTING↗

Chapter 14 - Reliability of Wind Turbines

The global wind energy industry has grown at a fast pace during the past half-decade. Advancements from design and manufacturing to operation and maintenance have led to reduced capital and maintenance costs, which make wind power an indispensable source for a comprehensive solution to global electricity needs. Once wind turbines are installed, the opportunity to lower wind power costs is mainly through improved operation and maintenance practices. Modern wind turbines are equipped with tens or hundreds of measurement channels and are generating an abundance of data, with lot of efforts being put into data analysis by both the research community and the industry. One type of analysis is through the exploration of reliability engineering methods based on readily available data or maintenance records collected at typical wind power plants. If adopted and conducted appropriately, these analyses can quickly save operation and maintenance costs in a potentially impactful manner. The wind industry has adopted this discipline more broadly in recent years. This chapter discusses wind turbine reliability by highlighting the methodology of reliability engineering life data analysis. It first briefly discusses the fundamentals of wind turbine reliability and the current industry status. Then, the reliability engineering method for life analysis, including data collection, model development, and forecasting, is presented in detail and illustrated through two case studies. The chapter concludes with some remarks on potential opportunities to improve wind turbine reliability. An owner and operator's perspective is taken and mechanical components are used to exemplify the potential benefits of reliability engineering analysis to improve wind turbine reliability and availability.

database↗

Data-Driven Operator Theoretic Methods for Phase Space Learning and Analysis

This paper uses data-driven operator theoretic approaches to explore the global phase space of a dynamical system. In this work, we defined conditions for discovering new invariant subspaces in the state space of a dynamical system starting from an invariant subspace based on the spectral properties of the Koopman operator. When the system evolution is known locally in several invariant subspaces in the state space of a dynamical system, a phase space stitching result is derived that yields the global Koopman operator. Additionally, in the case of equivariant systems, a phase space stitching result is developed to identify the global Koopman operator using the symmetry properties between the invariant subspaces of the dynamical system and time-series data from any one of the invariant subspaces. Finally, these results are extended to topologically conjugate dynamical systems; in particular, the relation between the Koopman tuple of topologically conjugate systems is established. The proposed results are demonstrated on several second-order nonlinear dynamical systems including a bistable toggle switch. Our method elucidates a strategy for designing discovery experiments: experiment execution can be done in many steps, and models from different invariant subspaces can be combined to approximate the global Koopman operator.

42 ENGINEERING↗

Holistic energy analysis method for thermal management architectures of data centers

Modern high-performance computing (HPC) data centers (DCs), particularly those supporting energy-intensive artificial intelligence (AI) workloads, face escalating thermal management challenges that degrade performance through thermal throttling and drive up cooling power consumption and operational costs. To address this challenge, many have developed a wide variety of thermal management solutions (single-phase, two-phase, direct, indirect, hybrid, and more) which attempt to cool HPC DCs effectively while attempting to minimize overall system power consumption. However, the analysis of these solutions and methods to effectively compare one with another is lacking. Overall power usage effectiveness (PUE) and total-power usage effectiveness (TUE) provide a metric to quantify power consumption but fail to identify components in the system which require further optimization. To address this, we propose a holistic analytical framework – the waterfall diagram (WFD) – which leverages a waterfall chart methodology, offering a comprehensive visualization of both the thermal management system loop and heat flow pathways from individual server components to the outdoor ambient. Use of the WFD enables graphical estimations of power efficiency and cooling performance across each component of a DC cooling system and complements Sankey-style energy flow visualizations by additionally resolving stage-wise temperature changes and incremental TUE contributions. The framework is used in conjunction with simulation-based approaches, to conduct a detailed pressure drop and flow distribution analysis aimed at identifying the optimal coolant distribution architecture for a single-phase direct-to-chip water-cooled DC, which serves as the baseline for subsequent WFD analysis. Among the evaluated architectures, the 3 U modular coolant distribution architecture is found to demonstrate the best performance, considering minimal pressure drop and uniform flow distribution. In addition, TUE is calculated for each cooling loop component based on its associated pressure drop and corresponding pumping power, which are integrated into the WFD. This correlation between TUE and local temperature offers immediate insight into the power efficiency and thermal performance contributions of individual components, facilitating further development and optimization. Examples of WFD applications are presented under varying thermal loads and ambient conditions, demonstrating reasonable cooling strategies. Notably, the 3 U modular architecture maintains a consistent chip case temperature of 85°C, achieving a TUE of 1.016 at ambient temperature of 47°C, and a TUE of 1.026 at ambient temperature of 52°C. The WFD methodology provides an efficient, holistic, and streamlined framework for DC thermal management architecture assessment and enables design optimization which is important for addressing the thermal-fluidic energy challenges of current and next-generation DCs.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Extraction of branching ratios from HERFD data

A quantitative analysis method has been developed that allows the cross calibration of separate Uranium M 4 and M 5 X-ray absorption spectra (XAS), in particular those collected with the new High Energy Resolution Fluorescence Detection (HERFD) method. With this method, it is now possible to generate experimental Branching Ratio (BR) values from the U M 4,5 XAS HERFD data.

36 MATERIALS SCIENCE↗

Measurement of Rayleigh Scattering in Liquid Argon at Vacuum Ultraviolet Wavelengths

Liquid argon based detectors are limited in scintillation light analysis capabilities due to inconsistent measurements of fundamental constants essential for the reconstruction of scintillation events. This experiment measures the Rayleigh scattering length of vacuum ultraviolet light propagating in liquid argon in support of liquid argon experiments. Preliminary results are presented for the Rayleigh scattering length in a range of vacuum ultraviolet wavelengths. Current data collection and analysis methods in this measurement require further advancements to understand and eliminate anomalous phenomena in the data. Future measurements in this experiment aim to precisely measure the liquid argon ultraviolet light scattering length. Results will contribute to the development of new photon system analysis methods for liquid argon experiments.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Improved Data Interpretation through Identification of Time Series Periodicity Changes

Analysis and interpretation of time series data is easiest when the data values occur at uniform intervals in time, but actual data may have differing data sampling frequencies, such as monthly and daily readings. Applying data analysis techniques, such as smoothing, to such a data set may not give a representative result between time segments. The ability to automatically distinguish time segments of differing data frequency would provide a means for applying data analysis independently to each segment, though a suitable blending at segment boundaries would be required. A method for detecting frequency changes was developed and applied to Gaussian and median smoothing of hydraulic head data from groundwater wells at the U.S. Department of Energy Hanford Site in southeastern Washington state. The process identifies time segments of high-frequency (daily) or low-frequency (greater than daily) data using adjusted-bandwidth Gaussian kernel density estimation and a threshold value, which are further refined to address small blocks of low-frequency data within larger blocks of high-frequency data. User-selectable levels of smoothing are then applied independently to the time segments prior to combining the segment results for a single smoothed data set. This time segment identification approach provides effective low- and high-frequency data separation, which provides a method to apply data analysis independently to each time segment.

97 MATHEMATICS AND COMPUTING↗

Methods for a blind analysis of isobar data collected by the STAR collaboration

In 2018, the STAR collaboration collected data from $_{44}^{96}{\mathrm{Ru}}+_{44}^{96}{\mathrm{Ru}}$ and $_{40}^{96}{\mathrm{Zr}}+_{40}^{96}{\mathrm{Zr}}$ at $\sqrt{s_\text {NN}}=200$ GeV to search for the presence of the chiral magnetic effect in collisions of nuclei. The isobar collision species alternated frequently between $_{44}^{96}{\mathrm{Ru}}+_{44}^{96}{\mathrm{Ru}}$ and $_{40}^{96}{\mathrm{Zr}}+_{40}^{96}{\mathrm{Zr}}$ . In order to conduct blind analyses of studies related to the chiral magnetic effect in these isobar data, STAR developed a three-step blind analysis procedure. Analysts are initially provided a “reference sample” of data, comprised of a mix of events from the two species, the order of which respects time-dependent changes in run conditions. After tuning analysis codes and performing time-dependent quality assurance on the reference sample, analysts are provided a species-blind sample suitable for calculating efficiencies and corrections for individual $\approx 30$-min data-taking runs. For this sample, species-specific information is disguised, but individual output files contain data from a single isobar species. Only run-by-run corrections and code alteration subsequent to these corrections are allowed at this stage. Following these modifications, the “frozen” code is passed over the fully un-blind data, completing the blind analysis. As a check of the feasibility of the blind analysis procedure, analysts completed a “mock data challenge,” analyzing data from Au + Au collisions at $\sqrt{s_\text {NN}}=27$ GeV, collected in 2018. The Au + Au data were prepared in the same manner intended for the isobar blind data. Finally, the details of the blind analysis procedure and results from the mock data challenge are presented.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Quantitative texture analysis using the NOMAD time-of-flight neutron diffractometer

Strategies for efficient and reliable texture measurements have been explored using the Nanoscale Ordered Materials Diffractometer (NOMAD) at the Spallation Neutron Source located at Oak Ridge National Laboratory (ORNL). To test these strategies, the texture of an Al alloy was also investigated using another neutron diffraction instrument, a constant-wavelength neutron diffractometer (NRSF2) located at the High Flux Isotope Reactor, also at ORNL. Reasonable agreement was found across the two experimental methods, but differences in overall texture strength and the symmetry of some components were noted, depending on the data reduction and analysis method selected. Finally, on the basis of these results, potential improvements are identified which would enhance the texture measurement capability on NOMAD.

47 OTHER INSTRUMENTATION↗

Preliminary Study on Fine-Grained Power and Energy Measurements on Grace Hopper GH200 with Open-Source Performance Tools

The increasing adoption of tightly integrated, heterogeneous architectures, combined with the slowdown of Moore’s law, has made application power and energy-driven optimizations critical to efficiently use high-performance computing systems. This paper introduces a newly developed open-source toolkit that seamlessly integrates the Linux real-time hardware monitoring program hwmon with the Performance Application Programming Interface and the Score-P performance measurement system, thereby enabling fine-grained power and energy measurements for high-performance computing applications. Our primary target platform is the Wombat test bed, which is a system based on the NVIDIA GH200 superchip. The toolkit can capture transient power peaks with high temporal resolution (50 ms) and, thanks to Score-P integration, can map power metrics to specific code regions, thereby providing actionable information on power-intensive operations and inefficiencies. The toolkit also provides a holistic view of both the power and the energy consumption of the entire GH200 superchip by covering all major components: the Grace CPU, the Hopper GPU, and the I/O subsystem. Experiments that use Locally Self-consistent Multiple Scattering, which is an application for first-principles calculations of materials developed at Oak Ridge National Laboratory, have demonstrated the tool’s ability to identify transient power spikes and uncover opportunities for energy-aware optimizations. Additionally, we introduce a Python-based utility for converting Open Trace Format 2 traces to Parquet format, thus enabling advanced data analysis for numerical integration methods applied to power data for accurate energy profiling.

Hernandez Mendoza, Oscar [ORNL] (ORCID:00000002538↗

Simulated JWST Data Sets for Multispectral and Hyperspectral Image Fusion

The James Webb Space Telescope (JWST) will provide multispectral and hyperspectral infrared images of a large number of astrophysical scenes. Multispectral images will have the highest angular resolution, while hyperspectral images (e.g., with integral field unit spectrometers) will provide the best spectral resolution. This paper aims at providing a comprehensive framework to generate an astrophysical scene and to simulate realistic hyperspectral and multispectral data acquired by two JWST instruments, namely, NIRCam Imager and NIRSpec IFU. We want to show that this simulation framework can be resorted to assess the benefits of fusing these images to recover an image of high spatial and spectral resolutions. To do so, we make a synthetic scene associated with a canonical infrared source, the Orion Bar. We develop forward models including corresponding noises for the two JWST instruments based on their physical features. JWST observations are then simulated by applying the forward models to the aforementioned synthetic scene. We test a dedicated fusion algorithm we developed on these simulated observations. We show that the fusion process reconstructs the high spatio-spectral resolution scene with a good accuracy on most areas, and we identify some limitations of the method to be tackled in future works. The synthetic scene and observations presented in the paper can be used, for instance, to evaluate instrument models, pipelines, or more sophisticated algorithms dedicated to JWST data analysis. Besides, fusion methods such as the one presented in this paper are shown to be promising tools to fully exploit the unprecedented capabilities of the JWST.

79 ASTRONOMY AND ASTROPHYSICS↗

Author Correction: US oil and gas system emissions from nearly one million aerial site measurements

Correction to: Naturehttps://doi.org/10.1038/s41586-024-07117-5 Published online 13 March 2024 In the version of the article initially published, several errors were present and have been corrected in the HTML and PDF versions of the article and Supplementary Information. The main results, conclusions, and our interpretations of the data remain unchanged. See the new Supplementary Information Section S15 for a more detailed description of the errors corrected and the resulting effects on the analysis. Data processing and methods corrections Overflight count correction: We previously used pre-computed source coverage data for some Carbon Mapper campaigns that was computed differently than was required for our analysis. We have re-computed Carbon Mapper source coverage based on flightline polygons and source coordinates. Transition point computation, well sites: The updated version now correctly compares the cumulative emissions distribution of simulated well site emissions with that of aerially detected sources (rather than plumes) when computing the transition point. Transition point computation, midstream: Additionally, the transition point calculation has been corrected to exclude aerially detected midstream emissions below the transition point, which was previously leading to double counting of these emissions. This error was not present for upstream (well site) emissions. Calculation errors Unit error: We corrected a specific unit conversion error affecting well site emissions in the Kairos Fort Worth dataset. Across all datasets, we also correct the conversion factor for converting from standard volume to mass for midstream emissions. Sorting error: We correct code that was applying incorrect sorting when computing correction factors to account for partial detection at well sites. Small typographical corrections were made in Fig. 1b and SI Section S4.1. Data processing and methods corrections Overflight count correction: We previously used pre-computed source coverage data for some Carbon Mapper campaigns that was computed differently than was required for our analysis. We have re-computed Carbon Mapper source coverage based on flightline polygons and source coordinates. Transition point computation, well sites: The updated version now correctly compares the cumulative emissions distribution of simulated well site emissions with that of aerially detected sources (rather than plumes) when computing the transition point. Transition point computation, midstream: Additionally, the transition point calculation has been corrected to exclude aerially detected midstream emissions below the transition point, which was previously leading to double counting of these emissions. This error was not present for upstream (well site) emissions. Calculation errors Unit error: We corrected a specific unit conversion error affecting well site emissions in the Kairos Fort Worth dataset. Across all datasets, we also correct the conversion factor for converting from standard volume to mass for midstream emissions. Sorting error: We correct code that was applying incorrect sorting when computing correction factors to account for partial detection at well sites. Small typographical corrections were made in Fig. 1b and SI Section S4.1. The following practices may help researchers conducting similar analyses avoid making similar errors: 1, Clear, accessible documentation explaining the interpretation of all columns in data input tables and all internal variables within the model, 2, Simple cross-check calculations computed before and after unit conversions.

Sherwin, Evan D↗

Electron Tomography and Machine Learning for Understanding the Highly Ordered Structure of Leafhopper Brochosomes

Insects known as leafhoppers (Hemiptera: Cicadellidae) produce hierarchically structured nanoparticles known as brochosomes that are exuded and applied to the insect cuticle, thereby providing camouflage and anti-wetting properties to aid insect survival. Although the physical properties of brochosomes are thought to depend on the leafhopper species, the structure–function relationships governing brochosome behavior are not fully understood. Brochosomes have complex hierarchical structures and morphological heterogeneity across species, due to which a multimodal characterization approach is required to effectively elucidate their nanoscale structure and properties. In this work, we study the structural and mechanical properties of brochosomes using a combination of atomic force microscopy (AFM), electron microscopy (EM), electron tomography, and machine learning (ML)-based quantification of large and complex scanning electron microscopy (SEM) image data sets. This suite of techniques allows for the characterization of internal and external brochosome structures, and ML-based image analysis methods of large data sets reveal correlations in the structure across several leafhopper species. Our results show that brochosomes are relatively rigid hollow spheres with characteristic dimensions and morphologies that depend on leafhopper species. Nanomechanical mapping AFM is used to determine a characteristic compression modulus for brochosomes on the order of 1–3 GPa, which is consistent with crystalline proteins. Altogether, this work provides an improved understanding of the structural and mechanical properties of leafhopper brochosomes using a new set of ML-based image classification tools that can be broadly applied to nanostructured biological materials.

Chemical structure↗