Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “large data sets”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Adaptive Machine Learning for Robust Diagnostics and Control of Time-Varying Particle Accelerator Components and Beams

Machine learning (ML) is growing in popularity for various particle accelerator applications including anomaly detection such as faulty beam position monitor or RF fault identification, for non-invasive diagnostics, and for creating surrogate models. ML methods such as neural networks (NN) are useful because they can learn input-output relationships in large complex systems based on large data sets. Once they are trained, methods such as NNs give instant predictions of complex phenomenon, which makes their use as surrogate models especially appealing for speeding up large parameter space searches which otherwise require computationally expensive simulations. However, quickly time varying systems are challenging for ML-based approaches because the actual system dynamics quickly drifts away from the description provided by any fixed data set, degrading the predictive power of any ML method, and limits their applicability for real time feedback control of quickly time-varying accelerator components and beams. In contrast to ML methods, adaptive model-independent feedback algorithms are by design robust to un-modeled changes and disturbances in dynamic systems, but are usually local in nature and susceptible to local extrema. In this work, we propose that the combination of adaptive feedback and machine learning, adaptive machine learning (AML), is a way to combine the global feature learning power of ML methods such as deep neural networks with the robustness of model-independent control. We present an overview of several ML and adaptive control methods, their strengths and limitations, and an overview of AML approaches.

97 MATHEMATICS AND COMPUTING↗

Training Ultrasound Image Classification Deep-Learning Algorithms for Pneumothorax Detection Using a Synthetic Tissue Phantom Apparatus

Ultrasound (US) imaging is a critical tool in emergency and military medicine because of its portability and immediate nature. However, proper image interpretation requires skill, limiting its utility in remote applications for conditions such as pneumothorax (PTX) which requires rapid intervention. Artificial intelligence has the potential to automate ultrasound image analysis for various pathophysiological conditions. Training models require large data sets and a means of troubleshooting in real-time for ultrasound integration deployment, and they also require large animal models or clinical testing. Here, we detail the development of a dynamic synthetic tissue phantom model for PTX and its use in training image classification algorithms. The model comprises a synthetic gelatin phantom cast in a custom 3D-printed rib mold and a lung mimicking phantom. When compared to PTX images acquired in swine, images from the phantom were similar in both PTX negative and positive mimicking scenarios. We then used a deep learning image classification algorithm, which we previously developed for shrapnel detection, to accurately predict the presence of PTX in swine images by only training on phantom image sets, highlighting the utility for a tissue phantom for AI applications.

Boice, Emily N. (ORCID:0000000171802842)↗

Learning to Count Grave Sites for Cemetery Observation Models With Satellite Imagery

Understanding how people occupy open spaces is important for research in support of population modeling, policy, national security, emergency response, and sustainability. For the past decade, there has been an increase in research toward capturing and reporting population dynamics and patterns of life at the building level and in some open public spaces such as cemeteries and parks. This is done through observation models developed from local sociocultural information acquired at various spatiotemporal scales to inform night, day, and episodic population occupancy estimates (people/1000 sq ft). Sociocultural information for cemeteries and parks is scarcely available and often collected manually. The process is not only marred by inconsistencies but is laborious and time consuming. In this study, we leverage convolutional neural networks (CNNs) and satellite imagery to derive grave site counts as proxy variables to support scalable and accurate sociocultural data required in a population observation model. Through a hybrid workflow (weak localization plus regression model), we characterize a large scale automation process to counting of grave sites. We evaluate and demonstrate the efficacy of proposed workflow using out-of-data set large satellite imagery and establish its broader impact on cemetery observation models.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Getting to the Core of PARAFAC2, A Nonnegative Approach

In this paper, the authors present a novel method of performing PARAFAC2 factorization of three-way data using a compact representation of that data. In the standard PARAFAC2 algorithm, two modes of the data are recovered directly during the decomposition while the third mode is returned as a transformation matrix, which is then used to rotate sets of orthogonal third-mode basis factors into interpretable factors. In our new method, the data are first decomposed into a core matrix and orthogonal factor loading matrices in the first two modes as well as sets of orthogonal factors in the third mode (as in standard PARAFAC2). The core matrix is then decomposed using a the standard PARAFAC2 strategy to produce transformation matrices in all three modes. The algorithm is particularly useful for very large data sets and essentially permits imposition of nonnegativity in all three modes.

97 MATHEMATICS AND COMPUTING↗

A geospatial risk analysis graphical user interface for identifying hazardous chemical emission sources

Background: Performing back trajectory and forward trajectory using the Hybrid Single-Particle Lagrangian Integrated Trajectory Model (HYSPLIT) is a reliable approach for assessing particle transport after release among mid-field atmospheric models. HYSPLIT has an externally facing online interface that allows non-expert users to run the model trajectories without requiring extensive training or programming. However, the existing HYSPLIT interface is limited if simulations have a large amount of meteorological data and timesteps that are not coincident. The objective of this study is to design and develop a more robust tool to rapidly evaluate hazard transport conditions and to perform risk analysis, while still maintaining an intuitive and user-friendly interface. Methods: HYSPLIT calculates forward and backward trajectories of particles based on wind speed, wind direction, and the corresponding location, timestamp, and Pasquill stability classes of the regions of the atmosphere in terms of the wind speed, the amount of solar radiation, and the fractional cloud cover. The computed particle transport trajectories, combined with the online Proton Transfer Reaction-Mass Spectrometry (PTR-MS) data (https://figshare.com/articles/dataset/ARL_Data_from_PROS_station_at_Hanford_site/19993964), can be used to identify and quantify the sources and affected area of the hazardous chemicals’ emission using the potential source distribution function (PSDF). PSDF is an improved statistical function based on the well-known potential source contribution function (PSCF) in establishing the air pollutant source and receptor relationship. Performing this analysis requires a range of meteorological and pollutant concentration measurements to be statistically meaningful. The existing HYSPLIT graphical user interface (GUI) does not easily permit computations of trajectories of a dataset of meteorological data in high temporal frequency. To improve the performance of HYSPLIT computations from a large dataset and enhance risk analysis of the accidental release of material at risk, a geospatial risk analysis tool (GRAT-GUI) is created to allow large data sets to be processed instantaneously and to provide ease of visualization. Results: The GRAT-GUI is a native desktop-based application and can be run in any Windows 10 system without any internet access requirements, thus providing a secure way to process large meteorological datasets even on a standalone computer. GRAT-GUI has features to import, integrate, and convert meteorological data with various formats for hazardous chemical emission source identification and risk analysis as a self-explanatory user interface. The tool is available at https://figshare.com/articles/software/GRAT/19426742.

97 MATHEMATICS AND COMPUTING↗

A new Monte Carlo generator for BSM physics in B → K*ℓ+ℓ− decays with an application to lepton non-universality in angular distributions

Abstract Within the widely used EvtGen framework, we have added a new event generator model forB → K * ℓ + ℓ − with improved standard model (SM) decay amplitudes and possible BSM physics contributions, which are implemented in the operator product expansion in terms of Wilson coefficients. This event generator can then be used to estimate the statistical sensitivity of a simulated experiment to the most general BSM signal resulting from dimension-six operators. We describe the advantages and potential of the newly developed ‘Sibidanov Physics Generator’ in improving the experimental sensitivity of searches for lepton non-universal BSM physics and clarifying signatures. The new generator can properly simulate BSM scenarios, interference between SM and BSM amplitudes, and correlations between different BSM observables as well as acceptance bias. We show that exploiting such correlations substantially improves experimental sensitivity. As a demonstration of the utility of the MC generator, we examine the prospects for improved measurements of lepton non-universality in angular distributions forB→K * ℓ + ℓ − decays from the expected 50 ab −1 data set of the Belle II experiment, using a four-dimensional unbinned maximum likelihood fit. We describe promising experimental signatures and correlations between observables. The use of lepton-universality violating ∆-observables significantly reduces uncertainties in the SM expectations due to QCD and resonance effects and is ideally suited for Belle II with the large data sets expected in the next decade. Thanks to the clean experimental environment of ane + e − machine, Belle II should be able to probe BSM physics in the Wilson coefficientsC 7 and$$ {C}_7^{\prime } $$ C 7 ′ , which appear at lowq 2 in the di-electron channel.

Physics↗

ParChain: a framework for parallel hierarchical agglomerative clustering using nearest-neighbor chain

This paper studies the hierarchical clustering problem, where the goal is to produce a dendrogram that represents clusters at varying scales of a data set. We propose the ParChain framework for designing parallel hierarchical agglomerative clustering (HAC) algorithms, and using the framework we obtain novel parallel algorithms for the complete linkage, average linkage, and Ward's linkage criteria. Compared to most previous parallel HAC algorithms, which require quadratic memory, our new algorithms require only linear memory, and are scalable to large data sets. ParChain is based on our parallelization of the nearest-neighbor chain algorithm, and enables multiple clusters to be merged on every round. We introduce two key optimizations that are critical for efficiency: a range query optimization that reduces the number of distance computations required when finding nearest neighbors of clusters, and a caching optimization that stores a subset of previously computed distances, which are likely to be reused. Experimentally, we show that our highly-optimized implementations using 48 cores with two-way hyper-threading achieve 5.8--110.1x speedup over state-of-the-art parallel HAC algorithms and achieve 13.75--54.23x self-relative speedup. Compared to state-of-the-art algorithms, our algorithms require up to 237.3x less space. Our algorithms are able to scale to data set sizes with tens of millions of points, which existing algorithms are not able to handle.

Computer Science↗

The scientific potential and technological challenges of the High-Luminosity Large Hadron Collider program

Here, we present an overview of the High-Luminosity (HL-LHC) program at the Large Hadron Collider (LHC), its scientific potential and technological challenges for both the accelerator and detectors. The HL-LHC program is expected to start circa 2027 and aims to increase the integrated luminosity delivered by the LHC by an order of magnitude at the collision energy of 14 TeV. This requires upgrades to the injector system, accelerator complex and luminosity levelling. The two experiments, ATLAS and CMS, require substantial upgrades to most of their systems in order to cope with the increased interaction rate, and much higher radiation levels than at the current LHC. We present selected examples based on novel ideas and technologies for applications at a hadron collider. Both experiments will replace their tracking systems. We describe the ATLAS pixel detector upgrade featuring novel tilted modules, and the CMS Outer Tracker upgrade with a new module design enabling use of tracks in the level-1 trigger system. CMS will also install state-of-the-art highly segmented calorimeter endcaps. Finally, we describe new picosecond precision timing detectors of both experiments. In addition, we discuss how the upgrades will enhance the physics performance of the experiments, and solve the computing challenges posed by the expected large data sets. The physics program of the HL-LHC is focused on precision measurements probing the limits of the Standard Model (SM) of particle physics and discovering new physics. We present a selection of studies that have been carried out to motivate the HL-LHC program. A central topic of exploration will be the characterization of the Higgs boson. The large HL-LHC data samples will extend the sensitivity of searches for new particles or new interactions whose existence has been hypothesized in order to explain shortcomings of the SM. Finally, we comment on the nature of large scientific collaborations.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Machine Learning Approaches to Predicting Induced Seismicity and Imaging Geothermal Reservoir Properties

This project developed machine learning (ML) methods, lab data sets, and field data to advance geothermal exploration and geothermal energy production. The work had three focus areas. One involved the development of ML methods to use microearthquakes (MEQs) for imaging geothermal reservoir properties and improving subsurface characterization – most importantly the evolution of permeability within the evolving reservoir. This part of the work included development of ML approaches for automated MEQ location, focal mechanism determination and identification of earthquake precursors. The second area focused on using MEQ signals generated by geothermal exploration and production to predict the relationship between fluid injection and seismicity. Here, we extended to reservoir scale our success in using ML to predict laboratory earthquakes and fault zone stress state. The third focus area was on lab experiments. Here, we developed new ML models for lab earthquake prediction and identification of precursors to failure to improve earthquake forecasting and early warning in geothermal settings. Major outcomes of our work include ML models that learn from MEQ signals during geothermal exploration and production to predict induced seismicity. MEQs occur naturally in connection with drilling and energy production. We developed ML methods to use the seismic waves from these events to characterize the elastic, hydraulic and poromechanical properties of reservoirs. Our work illuminated fracture geometry and the evolution of fracture permeability by incorporating seismic coda wave analysis and ML methods to relate fluid injection and seismicity. We significantly expanded laboratory earthquake prediction to include methods that use both passive measurements of microearthquakes within the lab fault zones and also active source acoustic measurements of fault zone elastic properties. These methods can now predict fault zone stress state, time to failure and the magnitude of lab earthquakes. Our work showed that repetitive stick- slip failure events during frictional sliding (the lab equivalent of earthquakes) are preceded by a cascade of micro-failure events that radiate energy in a manner that foretells unstable failure – manifest as laboratory MEQs. We documented a mapping between fracture properties and statistical attributes of elastic radiation. We extended existing works to geothermal reservoir scale and developed ML methods to determine reservoir permeability, fracture properties, and their evolution during geothermal energy production. An attractive feature of ML algorithms is their ability to handle big datasets and reveal patterns and correlations that may remain invisible to conventional analyses. Our work connected data from field, laboratory and intermediate scales to study permeability, stress, strength, fracture stiffness and geometry. At the field scale we used data from the Newberry Volcano field site, UtahFORGE, EGS Collab, and also the Bedretto underground research lab in Switzerland. These data sets are bridging the gap between the lab scale, theory, and reservoir scale. Our work produced plain language summaries to improve public understanding of DOE research. We also developed openly distributed ML and seismicity datasets for use by all researchers and we published connections between induced seismicity in geothermal areas and reservoir properties including permeability, fracture properties, and stress state. Our models are designed for the large data sets of induced seismicity typically associated with geothermal sites. We produced labeled event catalogs and used them on geothermal data to assess how ML can facilitate geothermal production and exploration. All datasets are available on the GDR Productivity: The project produced 32 publications in peer reviewed journals (two are in review). It supported the work of 6 PhD students, 40 conference presentations, 6 keynote talks at national meetings, and mentoring and professional development for 4 postdoctoral fellows.

15 GEOTHERMAL ENERGY↗

Indra: a public computationally accessible suite of cosmological N -body simulations

ABSTRACT Indra is a suite of large-volume cosmological N-body simulations with the goal of providing excellent statistics of the large-scale features of the distribution of dark matter. Each of the 384 simulations is computed with the same cosmological parameters and different initial phases, with 10243 dark matter particles in a box of length 1 h−1 Gpc, 64 snapshots of particle data and halo catalogues, and 505 time-steps of the Fourier modes of the density field, amounting to almost a petabyte of data. All of the Indra data are immediately available for analysis via the SciServer science platform, which provides interactive and batch computing modes, personal data storage, and other hosted data sets such as the Millennium simulations and many astronomical surveys. We present the Indra simulations, describe the data products and how to access them, and measure ensemble averages, variances, and covariances of the matter power spectrum, the matter correlation function, and the halo mass function to demonstrate the types of computations that Indra enables. We hope that Indra will be both a resource for large-scale structure research and a demonstration of how to make very large data sets public and computationally accessible.

Falck, Bridget↗

A new phenomenological model to describe root-soil interactions based on percolation theory

In his paper on net primary productivity of terrestrial communities predicted from climatological data, Rosenzweig (1968) argued that variability in productivity is well accounted for by (evapo)-transpiration, and that water from transpiration is, on global scales, the most variable component in the photosynthesis reaction. The goal of this paper is to investigate whether variability in plant growth on local scales and within species is primarily related to transpiration under several scenarios including different terrain curvature, slope aspect, soil characteristics, and climate ranges. Here, we test the hypothesis that this relationship exists because root growth into the surface soil layers (0–2 m) tends to follow paths with minima in resistance, which in turn maximizes water flow and nutrient delivery rates that regulate growth. The set of all connected paths with individual pore-to-pore flow resistances less than a critical, percolating, value forms a cluster with mass fractal dimensionality, d f . We propose that roots follow paths through the 2D percolation cluster, defining the set of all optimal flow paths, making the 2D value of d f from percolation relevant to root fractal dimensionality. The tortuosity of such optimal paths as defined in percolation theory should then relate root length to root radial extent, linking the parameters of root tortuosity and plant productivity. Our analysis of large data sets across species implies that root radial extent and tree height are both proportional to cumulative transpiration until trees approached maximum height, and their growth rates are proportional to the transpiration rate, not to the moisture content. Local variations in tree height as functions of the variables investigated appear generally consistent with deduced variations in transpiration. Here this correlation is investigated more closely in the context of studies addressing individual tree species.

54 ENVIRONMENTAL SCIENCES↗

Large-scale scenarios of electric vehicle charging with a data-driven model of control

Transportation electrification is forecast to bring millions of new electric vehicles to roads worldwide this decade. Planning to support those vehicles depends on detailed scenarios of their electricity demand in both uncontrolled and controlled or smart charging scenarios. In this work, we present a novel modeling approach to enable rapid generation of demand estimates that represent the impact of controlled charging for large-scale scenarios with millions of individual drivers. To model the effect of load modulation control on aggregate charging profiles, we propose a novel machine learning approach that replaces traditional optimization approaches. We demonstrate its performance modeling workplace charging control under a range of electricity rate schedules, achieving small errors (2.5%–4.5%) while accelerating computations by more than 4000 times. To generate the uncontrolled charging demand for scenarios with residential, workplace, and public charging we use statistical representations of a large data set of real charging sessions. We demonstrate the methodology by generating diverse sets of scenarios for California's charging demand in 2030 which consider multiple charging segments and controls, each run locally in under 50 s. We further demonstrate support for rate design by modeling the large-scale impact of a new, custom rate schedule for workplace charging.

33 ADVANCED PROPULSION SYSTEMS↗

Integrating HPC, AI, and Workflows for Scientific Data Analysis: Report from Dagstuhl Seminar 23352

The Dagstuhl Seminar 23352, titled “Integrating HPC, AI, and Workflows for Scientific Data Analysis,” held from August 27 to September 1, 2023, was a significant event focusing on the synergy between High-Performance Computing (HPC), Artificial Intelligence (AI), and scientific workflow technologies. The seminar recognized that modern Big Data analysis in science rests on three pillars: workflow technologies for reproducibility and steering, AI and Machine Learning (ML) for versatile analysis, and HPC for handling large data sets. These elements, while crucial, have traditionally been researched separately, leading to gaps in their integration. The seminar aimed to bridge these gaps, acknowledging the challenges and opportunities at the intersection of these technologies. The event highlighted the complex interplay between HPC, workflows, and ML, noting how ML has increasingly been integrated into scientific workflows, thereby enhancing resource demands and bringing new requirements to HPC architectures, like support for GPUs and iterative computations. The seminar also addressed the challenges in adapting HPC for large-scale ML tasks, including in areas like deep learning, and the need for workflow systems to evolve to leverage ML in data analysis fully. Moreover, the seminar explored how ML could optimize scientific workflow systems and HPC operations, such as through improved scheduling and fault tolerance. A key focus was on identifying prestigious use cases of ML in HPC and understanding their unique, unmet requirements. The stochastic nature of ML and its impact on the reproducibility of data analysis on HPC systems was also a topic of discussion.

97 MATHEMATICS AND COMPUTING↗

QUOTAS: A New Research Platform for the Data-driven Discovery of Black Holes

We present QUOTAS, a novel research platform for the data-driven investigation of supermassive black hole (SMBH) populations. While SMBH data—observations and simulations—have grown in complexity and abundance, our computational environments and tools have not matured commensurately to exhaust opportunities for discovery. To explore the BH, host galaxy, and parent dark matter halo connection—in this pilot version—we assemble and colocate the high-redshift, z > 3 quasar population alongside simulated data at the same cosmic epochs. As a first demonstration of the utility of QUOTAS, we investigate correlations between observed Sloan Digital Sky Survey (SDSS) quasars and their hosts with those derived from simulations. Leveraging machine-learning algorithms (ML), to expand simulation volumes, we show that halo properties extracted from smaller dark-matter-only simulation boxes successfully replicate halo populations in larger boxes. Next, using the Illustris-TNG300 simulation that includes baryonic physics as the training set, we populate the larger LEGACY Expanse dark-matter-only box with quasars, and show that observed SDSS quasar occupation statistics are accurately replicated. First science results from QUOTAS comparing colocated observational and ML-trained simulated data at z3 are presented. QUOTAS demonstrates the power of ML, in analyzing and exploring large data sets, while also offering a unique opportunity to interrogate theoretical assumptions that underpin accretion and feedback models. QUOTAS and all related materials are publicly available at the Google Kaggle platform. (The full data set—observational data and simulation data—are available at: https://www.kaggle.com/ and the codes are available at:https://www.kaggle.com/datasets/quotasplatform/quotas)

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

In-flight positional and energy use data set of a DJI Matrice 100 quadcopter for small package delivery

Abstract We autonomously directed a small quadcopter package delivery Uncrewed Aerial Vehicle (UAV) or “drone” to take off, fly a specified route, and land for a total of 209 flights while varying a set of operational parameters. The vehicle was equipped with onboard sensors, including GPS, IMU, voltage and current sensors, and an ultrasonic anemometer, to collect high-resolution data on the inertial states, wind speed, and power consumption. Operational parameters, such as commanded ground speed, payload, and cruise altitude, were varied for each flight. This large data set has a total flight time of 10 hours and 45 minutes and was collected from April to October of 2019 covering a total distance of approximately 65 kilometers. The data collected were validated by comparing flights with similar operational parameters. We believe these data will be of great interest to the research and industrial communities, who can use the data to improve UAV designs, safety, and energy efficiency, as well as advance the physical understanding of in-flight operations for package delivery drones.

42 ENGINEERING↗

J-PLUS: Support vector regression to measure stellar parameters

Stellar parameters are among the most important characteristics in studies of stars which, in traditional methods, are based on atmosphere models. However, time, cost, and brightness limits restrain the efficiency of spectral observations. The Javalambre Photometric Local Universe Survey (J-PLUS) is an observational campaign that aims to obtain photometry in 12 bands. Owing to its characteristics, J-PLUS data have become a valuable resource for studies of stars. Machine learning provides powerful tools for efficiently analyzing large data sets, such as the one from J-PLUS, and enables us to expand the research domain to stellar parameters. The main goal of this study is to construct a support vector regression (SVR) algorithm to estimate stellar parameters of the stars in the first data release of the J-PLUS observational campaign. The training data for the parameter's regressions are featured with 12-waveband photometry from J-PLUS and are crossidentified with spectrum-based catalogs. These catalogs are from the Large Sky Area Multi-Object Fiber Spectroscopic Telescope, the Apache Point Observatory Galactic Evolution Experiment, and the Sloan Extension for Galactic Understanding and Exploration. We then label them with the stellar effective temperature, the surface gravity, and the metallicity. Ten percent of the sample is held out to apply a blind test. We develop a new method, a multi-model approach, in order to fully take into account, the uncertainties of both the magnitudes and the stellar parameters. The method utilizes more than 200 models to apply the uncertainty analysis. We present a catalog of 2 493 424 stars with the root mean square error of 160 K in the effective temperature regression, 0.35 in the surface gravity regression, and 0.25 in the metallicity regression. We also discuss the advantages of this multi-model approach and compare it to other machine-learning methods.

79 ASTRONOMY AND ASTROPHYSICS↗

EQ_phase_detection

The EQ_phase_detection software is designed to scan continuous daily waveforms to detect earthquake phase arrivals from local to regional (150 km) events. The detections are made with a deep learning encoder-decoder model. When the model detects an earthquake in the waveforms, a second model is implemented to classify the first arriving motions. Both deep learning models are trained with the Tensorflow package using publicly available benchmark data sets. The software input is a path to a directory that contains waveforms in mseed format and the associated response files in xml format. The output is a data table of time stamped detections, signal amplitude, signal-to-noise ratio, and softmax probability of the detection in a generic format applicable to post-processing association algorithms for event locations. Additionally, the p-wave and s-wave waveforms are saved in a data table for rapid access when producing improved locations using correlation-based techniques. The software is designed for multiprocessing with multiple GPU’s for rapid processing of large data sets. The configuration file provides flexibility in the trained models implemented and allows access to multiple models trained for different sampling rates or input dimensions. This is particularly useful for regions with multiple networks that do not have the same data parameters.

Johnson, Christopher↗

Molecular Modeling and Molecular Dynamics Simulation of a Packed and Intact Bacterial Microcompartment

Bacterial microcompartments (BMCs) are protein-bound organelles found in some bacteria which encapsulate enzymes for enhanced catalytic activity. These compartments spatially sequester enzymes within semipermeable shell proteins and are packed full of enzyme cargoes and metabolites as they fulfill their function. Coupling together recent SAXS and proteomics work, it is possible to develop molecular models for these microcompartments and interrogate enzyme and metabolite dynamics within. Our primary goal of this study is to quantify the permeability of metabolite glyceraldehyde-3-phosphate (G3P) and dihydroxyacetone phosphate (DHAP) across the BMC shell through classical molecular dynamics simulation. The Haliangium ochraceum model of BMC shell (PDB: 6MZX) was used to model an intact BMC of approximately 10 million atoms. Working at this scale presented its own challenges in managing large data sets, with multiple challenges and hardware advances discussed that facilitated this work. Over approximately 750 ns of aggregate simulation, we see multiple permeation events for these metabolites that were added at high concentration through the pores present within BMC shell tiles. When compared to independent permeability estimates for the same metabolites determined through replica exchange umbrella sampling simulations, the permeabilities varied by approximately 3 orders of magnitude. Regardless, the permeability coefficients for both G3P and DHAP are highly similar and very high, such that only very small concentration gradients can be maintained across the BMC shell between the cytosol and BMC interior. The large simulation systems also facilitated comparisons for molecular diffusivity in the crowded environment within the BMC shell. By our estimates, the viscosity within a packed BMC shell is at least 10-fold higher than it would be in neat solution and is the real driver for varying permeability estimates we obtained through simulation. These findings will be used as design inputs for future bioengineering efforts to make products from BMCs, highlighting how permeable BMC shells can be.

Diffusion↗