Engineering Papers⌕ Search

Engineering topics

Bai, Tao

Publications and source records attributed to Bai, Tao.

Age over sex: evaluating gut microbiota differences in healthy Chinese populations

Age and gender have been recognized as two pivotal covariates affecting the composition of the gut microbiota. However, their mediated variations in microbiota seem to be inconsistent across different countries and races. In this study, 613 individuals, whom we referred to as the “healthy” population, were selected from 1,018 volunteers through rigorous selection using 16S rRNA sequencing. Three enterotypes were identified, namely, Escherichia–Shigella , mixture ( Bacteroides and Faecalibacterium ), and Prevotella . Moreover, 11 covariates that explain the differences in microbiota were determined, with age being the predominant factor. Furthermore, age-related differences in alpha diversity, beta diversity, and core genera were observed in our cohort. Remarkably, after adjusting for 10 covariates other than age, abundant genera that differed between age groups were demonstrated. In contrast, minimal differences in alpha diversity, beta diversity, and differentially abundant genera were observed between male and female individuals. Furthermore, we also demonstrated the age trajectories of several well-known beneficial genera, lipopolysaccharide (LPS)-producing genera, and short-chain fatty acids (SCFAs)-producing genera. Overall, our study further elucidated the effects mediated by age and gender on microbiota differences, which are of significant importance for a comprehensive understanding of the gut microbiome spectrum in healthy individuals.

Wu, Jiacheng↗

Quantitative methods and modeling to assess COVID–19–interrupted in vivo pharmacokinetic bioequivalence studies with two reference batches

The coronavirus disease 2019 (COVID-19) has presented unprecedented challenges to the generic drug development, including interruptions in bioequivalence (BE) studies. Per guidance published by the US Food and Drug Administration (FDA) during the COVID-19 public health emergency, any protocol changes or alternative statistical analysis plan for COVID-19-interrupted BE study should be accompanied with adequate justifications and not lead to biased equivalence determination. In this study, we used a modeling and simulation approach to assess the potential impact of study outcomes when two different batches of a Reference Standard (RS) were to be used in an in vivo pharmacokinetic BE study due to the RS expiration during the COVID-19 pandemic. Simulations were performed with hypothetical drugs under two scenarios: (1) uninterrupted study using a single batch of an RS, and (2) interrupted study using two batches of an RS. The acceptability of BE outcomes was evaluated by comparing the results obtained from interrupted studies with those from uninterrupted studies. The simulation results demonstrated that using a conventional statistical approach to evaluate BE for COVID-19-interrupted studies may be acceptable based on the pooled data from two batches. An alternative statistical method which includes a “batch” effect to the mixed effects model may be used when a significant “batch” effect was found in interrupted four-way crossover studies. However, such alternative method is not applicable for interrupted two-way crossover studies. Overall, the simulated scenarios are only for demonstration purpose, the acceptability of BE outcomes for the COVID19-interrupted studies could be case-specific.

60 APPLIED LIFE SCIENCES↗

Accelerating geostatistical modeling using geostatistics-informed machine Learning

Ordinary Kriging (OK) is a popular geostatistical algorithm for spatial interpolation and estimation. The computational complexity of OK changes quadratically and cubically for memory and speed, respectively, given the number of data. Therefore, it is computationally intensive and also challenging to process a large set of data, especially in three-dimensional (3D) cases. This paper develops a geostatistics-informed machine learning (GIML) model to improve the efficiency of OK by reducing the number of points required to be estimated using OK. Specifically, only a very few of the unknown points are estimated by OK to get the weights and estimations, which are used as the training dataset. Moreover, the governing equations of OK are used to guide our proposed machine learning to better reproduce the spatial distributions. Our results show that the proposed GIML can reduce the computational time of OK by at least one order of magnitude. The effectiveness of the GIML is evaluated and compared using a 2D case. Furthermore, we demonstrate its efficiency and robustness by considering a different number of training samples on various 3D simulation grids.

58 GEOSCIENCES↗

Hybrid geological modeling: Combining machine learning and multiple-point statistics

Accurately modeling and constructing a geologically realistic subsurface model remains an outstanding problem as the morphology controls the flow behaviors. Particularly, one of the pattern-based methods, namely cross-correlation based simulation, has been proved to be an effective way to reconstruct a realistic model, at both small and large scales. However, conditioning to point data in the large-scale problems is still a crucial issue in these algorithms, since there is always a trade-off between the quality of the realizations and the degree of point data reproduction. Specifically, it is not practical to build a training image (TI) which includes all the possibilities and variabilities. Therefore, finding a pattern that can represent the point data and, at the same time, preserving the connectivities is difficult. This leads to producing highly-connected realizations with a significant mismatch or poor models with a reasonable degree of point data reproduction. To accurately reproduce the densely distributed hard data, pixel-based methods can also produce some unrealistic artifacts around the hard data. In this paper, to overcome this challenge, however, we use pattern-based methods as they often produce more disconnected geobodies when dealing with dense hard data, and proposed a hybrid algorithm using the pattern-based methods and convolutional neural network (CNN). The trained CNN model is utilized to improve the quality of conditioning to point data for the original realizations generated by the pattern-based algorithm. As such, the mismatch locations are identified, and the same regions are used in the training of CNN to mimic the procedure through which a missing region can be filled. To evaluate the performance of the proposed hybrid algorithm, it is tested on cases with different dimensions and different numbers of facies. Then, the newly improved realizations are compared with the initial realizations generated by the pattern-based algorithm. The comparison is also conducted by the flow simulation test. And it indicates that the proposed hybrid algorithm can better reproduce the point data, while the connectivities are better preserved.

58 GEOSCIENCES↗