Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “large data sets”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Indra: a public computationally accessible suite of cosmological N -body simulations

ABSTRACT Indra is a suite of large-volume cosmological N-body simulations with the goal of providing excellent statistics of the large-scale features of the distribution of dark matter. Each of the 384 simulations is computed with the same cosmological parameters and different initial phases, with 10243 dark matter particles in a box of length 1 h−1 Gpc, 64 snapshots of particle data and halo catalogues, and 505 time-steps of the Fourier modes of the density field, amounting to almost a petabyte of data. All of the Indra data are immediately available for analysis via the SciServer science platform, which provides interactive and batch computing modes, personal data storage, and other hosted data sets such as the Millennium simulations and many astronomical surveys. We present the Indra simulations, describe the data products and how to access them, and measure ensemble averages, variances, and covariances of the matter power spectrum, the matter correlation function, and the halo mass function to demonstrate the types of computations that Indra enables. We hope that Indra will be both a resource for large-scale structure research and a demonstration of how to make very large data sets public and computationally accessible.

Falck, Bridget↗

A new phenomenological model to describe root-soil interactions based on percolation theory

In his paper on net primary productivity of terrestrial communities predicted from climatological data, Rosenzweig (1968) argued that variability in productivity is well accounted for by (evapo)-transpiration, and that water from transpiration is, on global scales, the most variable component in the photosynthesis reaction. The goal of this paper is to investigate whether variability in plant growth on local scales and within species is primarily related to transpiration under several scenarios including different terrain curvature, slope aspect, soil characteristics, and climate ranges. Here, we test the hypothesis that this relationship exists because root growth into the surface soil layers (0–2 m) tends to follow paths with minima in resistance, which in turn maximizes water flow and nutrient delivery rates that regulate growth. The set of all connected paths with individual pore-to-pore flow resistances less than a critical, percolating, value forms a cluster with mass fractal dimensionality, d f . We propose that roots follow paths through the 2D percolation cluster, defining the set of all optimal flow paths, making the 2D value of d f from percolation relevant to root fractal dimensionality. The tortuosity of such optimal paths as defined in percolation theory should then relate root length to root radial extent, linking the parameters of root tortuosity and plant productivity. Our analysis of large data sets across species implies that root radial extent and tree height are both proportional to cumulative transpiration until trees approached maximum height, and their growth rates are proportional to the transpiration rate, not to the moisture content. Local variations in tree height as functions of the variables investigated appear generally consistent with deduced variations in transpiration. Here this correlation is investigated more closely in the context of studies addressing individual tree species.

54 ENVIRONMENTAL SCIENCES↗

The Cabauw Intercomparison Campaign for Nitrogen Dioxide Measuring Instruments (CINDI): Design, Execution, and Early Results

From June to July 2009 more than thirty different in-situ and remote sensing instruments from all over the world participated in the Cabauw Intercomparison campaign for Nitrogen Dioxide measuring Instruments (CINDI). The campaign took place at KNMI's Cabauw Experimental Site for Atmospheric Research (CESAR) in the Netherlands. Its main objectives were to determine the accuracy of state-ofthe- art ground-based measurement techniques for the detection of atmospheric nitrogen dioxide (both in-situ and remote sensing), and to investigate their usability in satellite data validation. The expected outcomes are recommendations regarding the operation and calibration of such instruments, retrieval settings, and observation strategies for the use in ground-based networks for air quality monitoring and satellite data validation. Twenty-four optical spectrometers participated in the campaign, of which twenty-one had the capability to scan different elevation angles consecutively, the so-called Multi-axis DOAS systems, thereby collecting vertical profile information, in particular for nitrogen dioxide and aerosol. Various in-situ samplers and lidar instruments simultaneously characterized the variability of atmospheric trace gases and the physical properties of aerosol particles. A large data set of continuous measurements of these atmospheric constituents has been collected under various meteorological conditions and air pollution levels. Together with the permanent measurement capability at the CESAR site characterizing the meteorological state of the atmosphere, the CINDI campaign provided a comprehensive observational data set of atmospheric constituents in a highly polluted region of the world during summertime. First detailed comparisons performed with the CINDI data show that slant column measurements of NO2, O4 and HCHO with MAX-DOAS agree within 5 to 15%, vertical profiles of NO2 derived from several independent instruments agree within 25% of one another, and MAX-DOAS aerosol optical thickness agrees within 20-30% with AERONET data. For the in-situ NO2 instrument using a molybdenum converter, a bias was found as large as 5 ppbv during day time, when compared to the other in-situ instruments using photolytic converters.

Atmospheric composition↗

An interdisciplinary analysis of multispectral satellite data for selected cover types in the Colorado Mountains, using automatic data processing techniques

The author has reported the following significant results. A data set containing SKYLAB, LANDSAT, and topographic data has been overlayed, registered, and geometrically corrected to a scale of 1:24,000. After geometrically correcting both sets of data, the SKYLAB data were overlayed on the LANDSAT data. Digital topographic data were then obtained, reformatted, and a data channel containing elevation information was then digitally overlayed onto the LANDSAT and SKYLAB spectral data. The 14,039 square kilometers involving 2,113, 776 LANDSAT pixels represents a relatively large data set available for digital analysis. The overlayed data set enables investigators to numerically analyze and compare two sources of spectral data and topographic data from any point in the scene. This capability is new and it will permit a numerical comparison of spectral response with elevation, slope, and aspect. Utilization of the spectral and topographic data together to obtain more accurate classifications of the various cover types present is feasible.

Hoffer, R. M.↗

Large-scale scenarios of electric vehicle charging with a data-driven model of control

Transportation electrification is forecast to bring millions of new electric vehicles to roads worldwide this decade. Planning to support those vehicles depends on detailed scenarios of their electricity demand in both uncontrolled and controlled or smart charging scenarios. In this work, we present a novel modeling approach to enable rapid generation of demand estimates that represent the impact of controlled charging for large-scale scenarios with millions of individual drivers. To model the effect of load modulation control on aggregate charging profiles, we propose a novel machine learning approach that replaces traditional optimization approaches. We demonstrate its performance modeling workplace charging control under a range of electricity rate schedules, achieving small errors (2.5%–4.5%) while accelerating computations by more than 4000 times. To generate the uncontrolled charging demand for scenarios with residential, workplace, and public charging we use statistical representations of a large data set of real charging sessions. We demonstrate the methodology by generating diverse sets of scenarios for California's charging demand in 2030 which consider multiple charging segments and controls, each run locally in under 50 s. We further demonstrate support for rate design by modeling the large-scale impact of a new, custom rate schedule for workplace charging.

33 ADVANCED PROPULSION SYSTEMS↗

Integrating HPC, AI, and Workflows for Scientific Data Analysis: Report from Dagstuhl Seminar 23352

The Dagstuhl Seminar 23352, titled “Integrating HPC, AI, and Workflows for Scientific Data Analysis,” held from August 27 to September 1, 2023, was a significant event focusing on the synergy between High-Performance Computing (HPC), Artificial Intelligence (AI), and scientific workflow technologies. The seminar recognized that modern Big Data analysis in science rests on three pillars: workflow technologies for reproducibility and steering, AI and Machine Learning (ML) for versatile analysis, and HPC for handling large data sets. These elements, while crucial, have traditionally been researched separately, leading to gaps in their integration. The seminar aimed to bridge these gaps, acknowledging the challenges and opportunities at the intersection of these technologies. The event highlighted the complex interplay between HPC, workflows, and ML, noting how ML has increasingly been integrated into scientific workflows, thereby enhancing resource demands and bringing new requirements to HPC architectures, like support for GPUs and iterative computations. The seminar also addressed the challenges in adapting HPC for large-scale ML tasks, including in areas like deep learning, and the need for workflow systems to evolve to leverage ML in data analysis fully. Moreover, the seminar explored how ML could optimize scientific workflow systems and HPC operations, such as through improved scheduling and fault tolerance. A key focus was on identifying prestigious use cases of ML in HPC and understanding their unique, unmet requirements. The stochastic nature of ML and its impact on the reproducibility of data analysis on HPC systems was also a topic of discussion.

97 MATHEMATICS AND COMPUTING↗

A Web Architecture to Geographically Interrogate CHIRPS Rainfall and eMODIS NDVI for Land Use Change

Monitoring of rainfall and vegetation over the continent of Africa is important for assessing the status of crop health and agriculture, along with long‐term changes in land use change. These issues can be addressed through examination of long‐term precipitation (rainfall) data sets and remote sensing of land surface vegetation and land use types. Two products have been used previously to address these goals: the Climate Hazard Group Infrared Precipitation with Stations (CHIRPS) rainfall data, and multi‐day composites of Normalized Difference Vegetation Index (NDVI) from the USGS eMODIS product. Combined, these are very large data sets that require unique tools and architecture to facilitate a variety of data analysis methods or data exploration by the end user community. To address these needs, a web‐enabled system has been developed to allow end‐users to interrogate CHIRPS rainfall and eMODIS NDVI data over the continent of Africa. The architecture allows end‐users to use custom defined geometries, or the use of predefined political boundaries in their interrogation of the data. The massive amount of data interrogated by the system allows the end‐users with only a web browser to extract vital information in order to investigate land use change and its causes. The system can be used to generate daily, monthly and yearly averages over a geographical area and range of dates of interest to the user. It also provides analysis of trends in precipitation or vegetation change for times of interest. The data provided back to the end‐user is displayed in graphical form and can be exported for use in other, external tools. The development of this tool has significantly decreased the investment and requirements for end‐users to use these two important datasets, while also allowing the flexibility to the end‐user to limit the search to the area of interest.

Burks, Jason E.↗

QUOTAS: A New Research Platform for the Data-driven Discovery of Black Holes

We present QUOTAS, a novel research platform for the data-driven investigation of supermassive black hole (SMBH) populations. While SMBH data—observations and simulations—have grown in complexity and abundance, our computational environments and tools have not matured commensurately to exhaust opportunities for discovery. To explore the BH, host galaxy, and parent dark matter halo connection—in this pilot version—we assemble and colocate the high-redshift, z > 3 quasar population alongside simulated data at the same cosmic epochs. As a first demonstration of the utility of QUOTAS, we investigate correlations between observed Sloan Digital Sky Survey (SDSS) quasars and their hosts with those derived from simulations. Leveraging machine-learning algorithms (ML), to expand simulation volumes, we show that halo properties extracted from smaller dark-matter-only simulation boxes successfully replicate halo populations in larger boxes. Next, using the Illustris-TNG300 simulation that includes baryonic physics as the training set, we populate the larger LEGACY Expanse dark-matter-only box with quasars, and show that observed SDSS quasar occupation statistics are accurately replicated. First science results from QUOTAS comparing colocated observational and ML-trained simulated data at z3 are presented. QUOTAS demonstrates the power of ML, in analyzing and exploring large data sets, while also offering a unique opportunity to interrogate theoretical assumptions that underpin accretion and feedback models. QUOTAS and all related materials are publicly available at the Google Kaggle platform. (The full data set—observational data and simulation data—are available at: https://www.kaggle.com/ and the codes are available at:https://www.kaggle.com/datasets/quotasplatform/quotas)

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

In-flight positional and energy use data set of a DJI Matrice 100 quadcopter for small package delivery

Abstract We autonomously directed a small quadcopter package delivery Uncrewed Aerial Vehicle (UAV) or “drone” to take off, fly a specified route, and land for a total of 209 flights while varying a set of operational parameters. The vehicle was equipped with onboard sensors, including GPS, IMU, voltage and current sensors, and an ultrasonic anemometer, to collect high-resolution data on the inertial states, wind speed, and power consumption. Operational parameters, such as commanded ground speed, payload, and cruise altitude, were varied for each flight. This large data set has a total flight time of 10 hours and 45 minutes and was collected from April to October of 2019 covering a total distance of approximately 65 kilometers. The data collected were validated by comparing flights with similar operational parameters. We believe these data will be of great interest to the research and industrial communities, who can use the data to improve UAV designs, safety, and energy efficiency, as well as advance the physical understanding of in-flight operations for package delivery drones.

42 ENGINEERING↗

Accelerated Simulation of Air Pollution Using NVIDIA RAPIDS

Atmospheric chemistry models are a central tool to study and forecast the impact of air pollution on the environment, vegetation, and human health. However, the numerical simulation of chemical kinetics is computationally expensive due to the stiffness of the system of ordinary differential equations that describes atmospheric chemistry. Here we present an alternative approach to the computation of atmospheric chemistry based on machine learning. Our training data set is produced using the NASA Goddard Earth Observing System (GEOS) model with GEOS-Chem chemistry, run on the NASA Center for Climate Simulation (NCCS) Discover supercomputing cluster on 384 Intel Xeon Haswell cores. This model spends more than 50% of total run time on solving atmospheric chemistry. The data set contains as input features the air pollution concentrations before solving the differential equations, together with some key physical parameters such as temperature and sun intensity. As target variables we define the air pollution concentrations after solving the differential equations. Using Dask-cuDF and Dask-XGBoost on the NVIDIA RAPIDS platform on 8 Tesla V100 GPUs, we generate from this training set gradient boosted decision tree models that can reproduce the simulation of chemical kinetics. We do this on the NCCS Advanced Data Analytics Platform (ADAPT) science cloud environment. Our application takes full advantage of recent advances in Dask-XGBoost, such as multi-node and multi-GPU scaling for distributed training with large data sets. The increase in training data size enabled by this is critical to capture the full range of chemical environments encountered across the globe and all annual seasons.The boosted tree models offer good predictability and show many of the features of the full chemistry reference simulation. Further improvements can be achieved through mass balance considerations and by accounting for error correlations. We incorporate the boosted tree models into the GEOS reference model using XGBoost's C API. This enables a seamless integration of the GPU trained models into GEOS-Chem, which is written in Fortran and optimized for use in a massively parallel CPU environment. We show the benefits of this approach and discuss the potential speedup of this machine learning accelerated atmospheric chemistry model.

Keller, Christoph A.↗

Caloris Basin - An enhanced source for potassium in Mercury's atmosphere

Enhanced abundances of neutral K in the atmosphere of Mercury have been found above the longitude range containing Caloris Basin. Results of a large data set including six elongations of the planet between June 1986 and January 1988 show typical K column abundances of about 5.4 x 10 to the 8th K atmos/sq cm. During the observing period in October 1987, when Caloris Basin was in view, the typical K column was about 2.7 x 10 to the 9th K atoms/sq cm. Another large value was seen over the Caloris antipode in January 1988. This enhancement is consistent with an increased source of K from the well-fractured crust and regolith associated with this large impact basin. The phenomenon is localized because at most solar angles, thermal alkali atoms cannot move more than a few hundred kilometers from their source before being lost to ionization by solar ultraviolet radiation.

Sprague, Ann L.↗

Visualization in the Design of Modern Aircraft Aerodynamics

Modem aircraft design involves study of airflow through both windtunnel testing and computer simulation. These computer simulations result in often very large and complex sets of numbers, which contain information critical to the aircrafts performance. This talk will describe how visualization is used to understand these simulations, using a variety of techniques including low-level analysis such as simulated particles, high-level feature detection, and virtual-reality-based techniques for exploration. We will focus on the challenges of extremely large data sets, interactive performance, and information extraction. The talk will close with a vision of the future including the integration of simulation and visualization.

Bryson, Steve↗

Automatic variable selection in ecological niche modeling: A case study using Cassin’s Sparrow (Peucaea cassinii)

MERRA/Max provides a feature selection approach to dimensionality reduction that enables direct use of global climate model outputs in ecological niche modeling. The system accomplishes this reduction through a Monte Carlo optimization in which many independent MaxEnt runs, operating on a species occurrence file and a small set of randomly selected variables in a large collection of variables, converge on an estimate of the top contributing predictors in the larger collection. These top predictors can be viewed as potential candidates in the variable selection step of the ecological niche modeling process. MERRA/Max’s Monte Carlo algorithm operates on files stored in the underlying filesystem, making it scalable to large data sets. Its software components can run as parallel processes in a high-performance cloud computing environment to yield near real-time performance. In tests using Cassin’s Sparrow (Peucaea cassinii) as the target species, MERRA/Max selected a set of predictors from Worldclim’s Bioclim collection of 19 environmental variables that have been shown to be important determinants of the species’ bioclimatic niche. It also selected biologically and ecologically plausible predictors from a more diverse set of 86 environmental variables derived from NASA’s Modern-Era Retrospective Analysis for Research and Applications Version 2 (MERRA-2) reanalysis, an output product of the Goddard Earth Observing System Version 5 (GEOS-5) modeling system. We believe these results point to a technological approach that could expand the use global climate model outputs in ecological niche modeling, foster exploratory experimentation with otherwise difficult-to-use climate data sets, streamline the modeling process, and, eventually, enable automated bioclimatic modeling as a practical, readily accessible, low-cost, commercial cloud service.

John L. Schnase↗

A Fast Implementation of the ISOCLUS Algorithm

Unsupervised clustering is a fundamental tool in numerous image processing and remote sensing applications. For example, unsupervised clustering is often used to obtain vegetation maps of an area of interest. This approach is useful when reliable training data are either scarce or expensive, and when relatively little a priori information about the data is available. Unsupervised clustering methods play a significant role in the pursuit of unsupervised classification. One of the most popular and widely used clustering schemes for remote sensing applications is the ISOCLUS algorithm, which is based on the ISODATA method. The algorithm is given a set of n data points (or samples) in d-dimensional space, an integer k indicating the initial number of clusters, and a number of additional parameters. The general goal is to compute a set of cluster centers in d-space. Although there is no specific optimization criterion, the algorithm is similar in spirit to the well known k-means clustering method in which the objective is to minimize the average squared distance of each point to its nearest center, called the average distortion. One significant feature of ISOCLUS over k-means is that clusters may be merged or split, and so the final number of clusters may be different from the number k supplied as part of the input. This algorithm will be described in later in this paper. The ISOCLUS algorithm can run very slowly, particularly on large data sets. Given its wide use in remote sensing, its efficient computation is an important goal. We have developed a fast implementation of the ISOCLUS algorithm. Our improvement is based on a recent acceleration to the k-means algorithm, the filtering algorithm, by Kanungo et al.. They showed that, by storing the data in a kd-tree, it was possible to significantly reduce the running time of k-means. We have adapted this method for the ISOCLUS algorithm. For technical reasons, which are explained later, it is necessary to make a minor modification to the ISOCLUS specification. We provide empirical evidence, on both synthetic and Landsat image data sets, that our algorithm's performance is essentially the same as that of ISOCLUS, but with significantly lower running times. We show that our algorithm runs from 3 to 30 times faster than a straightforward implementation of ISOCLUS. Our adaptation of the filtering algorithm involves the efficient computation of a number of cluster statistics that are needed for ISOCLUS, but not for k-means.

Memarsadeghi, Nargess↗

J-PLUS: Support vector regression to measure stellar parameters

Stellar parameters are among the most important characteristics in studies of stars which, in traditional methods, are based on atmosphere models. However, time, cost, and brightness limits restrain the efficiency of spectral observations. The Javalambre Photometric Local Universe Survey (J-PLUS) is an observational campaign that aims to obtain photometry in 12 bands. Owing to its characteristics, J-PLUS data have become a valuable resource for studies of stars. Machine learning provides powerful tools for efficiently analyzing large data sets, such as the one from J-PLUS, and enables us to expand the research domain to stellar parameters. The main goal of this study is to construct a support vector regression (SVR) algorithm to estimate stellar parameters of the stars in the first data release of the J-PLUS observational campaign. The training data for the parameter's regressions are featured with 12-waveband photometry from J-PLUS and are crossidentified with spectrum-based catalogs. These catalogs are from the Large Sky Area Multi-Object Fiber Spectroscopic Telescope, the Apache Point Observatory Galactic Evolution Experiment, and the Sloan Extension for Galactic Understanding and Exploration. We then label them with the stellar effective temperature, the surface gravity, and the metallicity. Ten percent of the sample is held out to apply a blind test. We develop a new method, a multi-model approach, in order to fully take into account, the uncertainties of both the magnitudes and the stellar parameters. The method utilizes more than 200 models to apply the uncertainty analysis. We present a catalog of 2 493 424 stars with the root mean square error of 160 K in the effective temperature regression, 0.35 in the surface gravity regression, and 0.25 in the metallicity regression. We also discuss the advantages of this multi-model approach and compare it to other machine-learning methods.

79 ASTRONOMY AND ASTROPHYSICS↗

Wavenumber-frequency Spectra of Pressure Fluctuations on a Generic Space Vehicle Measured via Unsteady Pressure-Sensitive Paint

Time histories of pressure fluctuations on a generic, hammerhead space vehicle model were measured using unsteady Pressure-Sensitive Paint (uPSP). The test was conducted in the 11-foot transonic wind tunnel of NASA Ames Research Center over a Mach number range of 0.6 M 1.2, and angles of attack of -4 4. The model was coated with a porous binder and PtTFPP-based porous polymer paint. An elaborate system of four high-speed cameras, and forty LED lamps was used for image acquisition. Various steps for image registration, reduction of shot noise, photogrammetry procedure to map images from the four cameras on a grid for the model, and finally a calibration procedure to convert the measured fluctuations in light intensity to fluctuating pressure, are discussed in the paper. The calibration process using a set of unsteady pressure sensors mounted on the model, was found to overcome some of the inherent problems of the fast response paint, such as rapid photo-degradation, non-linearity in pressure response, and significant temperature sensitivity. Comparison of spectra of pressure fluctuations between UPSP and pressure sensors demonstrated the ability of the paint to faithfully follow fluctuations up to 10 kHz, the maximum attempted. It was also found that the camera bit-depth and the illumination level limited the lowest measurable levels of pressure fluctuations to around 140dB. The large data set exposed various critical transonic flow physics not seen before, such as a coupling of the shock motion on the Payload Fairing (PF) with the separated flow region on the upper stage of the launch vehicle, and upstream convection of pressure fluctuation on PF at certain Mach numbers. The data also confirmed the expectation of a general lowering of the coefficient of pressure fluctuation with Mach number. The availability of the data set on a dense, regularly-spaced, surface grid allowed for the calculation of wavenumber-frequency (k-) spectra via straightforward applications of Fourier transform. The k- spectra were compared for the separated flow regions on the Second Stage, and the shock-boundary layer interactions on PF. The former showed self-similarity with Mach number while the latter was distinctly different, and confirmed the upstream propagation of pressure fluctuations. The k- spectra were dominated by the convected fluctuations; the acoustic domain was not discernable. These data, valuable for the vibro-acoustics analysis of aerospace vehicles, are believed to be the first obtained for the transonic flight regime, and pave the path for application on production models of aerospace vehicles.

buffet↗

EQ_phase_detection

The EQ_phase_detection software is designed to scan continuous daily waveforms to detect earthquake phase arrivals from local to regional (150 km) events. The detections are made with a deep learning encoder-decoder model. When the model detects an earthquake in the waveforms, a second model is implemented to classify the first arriving motions. Both deep learning models are trained with the Tensorflow package using publicly available benchmark data sets. The software input is a path to a directory that contains waveforms in mseed format and the associated response files in xml format. The output is a data table of time stamped detections, signal amplitude, signal-to-noise ratio, and softmax probability of the detection in a generic format applicable to post-processing association algorithms for event locations. Additionally, the p-wave and s-wave waveforms are saved in a data table for rapid access when producing improved locations using correlation-based techniques. The software is designed for multiprocessing with multiple GPU’s for rapid processing of large data sets. The configuration file provides flexibility in the trained models implemented and allows access to multiple models trained for different sampling rates or input dimensions. This is particularly useful for regions with multiple networks that do not have the same data parameters.

Johnson, Christopher↗

Molecular Modeling and Molecular Dynamics Simulation of a Packed and Intact Bacterial Microcompartment

Bacterial microcompartments (BMCs) are protein-bound organelles found in some bacteria which encapsulate enzymes for enhanced catalytic activity. These compartments spatially sequester enzymes within semipermeable shell proteins and are packed full of enzyme cargoes and metabolites as they fulfill their function. Coupling together recent SAXS and proteomics work, it is possible to develop molecular models for these microcompartments and interrogate enzyme and metabolite dynamics within. Our primary goal of this study is to quantify the permeability of metabolite glyceraldehyde-3-phosphate (G3P) and dihydroxyacetone phosphate (DHAP) across the BMC shell through classical molecular dynamics simulation. The Haliangium ochraceum model of BMC shell (PDB: 6MZX) was used to model an intact BMC of approximately 10 million atoms. Working at this scale presented its own challenges in managing large data sets, with multiple challenges and hardware advances discussed that facilitated this work. Over approximately 750 ns of aggregate simulation, we see multiple permeation events for these metabolites that were added at high concentration through the pores present within BMC shell tiles. When compared to independent permeability estimates for the same metabolites determined through replica exchange umbrella sampling simulations, the permeabilities varied by approximately 3 orders of magnitude. Regardless, the permeability coefficients for both G3P and DHAP are highly similar and very high, such that only very small concentration gradients can be maintained across the BMC shell between the cytosol and BMC interior. The large simulation systems also facilitated comparisons for molecular diffusivity in the crowded environment within the BMC shell. By our estimates, the viscosity within a packed BMC shell is at least 10-fold higher than it would be in neat solution and is the real driver for varying permeability estimates we obtained through simulation. These findings will be used as design inputs for future bioengineering efforts to make products from BMCs, highlighting how permeable BMC shells can be.

Diffusion↗