Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Mathematical Methods In Social Sciences”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

29 records · Page 2

Molecular Hypernetworks for Exploration of Multi-Dimensional Metabolomics Data (Chyper)

Orthogonal separations of data from high-resolution mass spectrometry can provide insight into sample composition and help address the challenge of complete annotation of molecules in untargeted metabolomics. “Molecular networks” (MNs), as used, for example, in the Global Natural Products Social Molecular Networking platform, are an increasingly popular computational strategy for exploring and visualizing molecular relationships and improving annotation. MNs use graph representations to show the relationships between measured multidimensional data features. MNs also show promise for using network science algorithms to automatically identify targets for annotation candidates and to dereplicate features associated to a single molecular identity. However, more advanced methods may better represent the complexity present in samples. Our work aims to increase confidence in annotation propagation by extending molecular network methods to include “molecular hypernetworks” (MHNs), able to natively represent multiway relationships among observations supporting both human and analytical processing. In this paper we first introduce MHNs illustrated with simple examples, and demonstrate how to build them from liquid chromatography- and ion mobility spectrometry- separated MS data. We then describe a method to construct MHNs directly from existing MNs as their “clique reconstructions”, demonstrating their utility by comparing examples of previously published graph-based MNs to their respective MHNs.

59 BASIC BIOLOGICAL SCIENCES↗

COVID-19: Spatiotemporal social data analytics and machine learning for pandemic exploration and forecasting

This task focused on developing a preliminary approach to use machine learning (ML) to explore the relationship between county-level societal variables and COVID-19 parameters, including COVID-19 cases rates and counts and COVID-19 death rates and counts. The objective was to develop and test a prototype approach for linking COVID-19 and county-level data. The task focused on enhancing and applying existing LANL ML techniques to COVID-19. Our novel ML methods have been a subject of a recently approved U.S. patent. The codes based on these methods are already open-source released. Our ML tools (NMFk/NTFk) are applied to extract hidden features (signals, waves) in the analyzed datasets and automatically identify their optimal number. The features are extracted by identifying counties that have similarities between the county-level societal variables and the COVID-19 parameters. These demonstration analyses will facilitate the ongoing pandemic simulations and predictions performed by Los Alamos other institutions, as well as lay the groundwork for future work.

60 APPLIED LIFE SCIENCES↗

Optimal experimental design: Formulations and computations

Questions of ‘how best to acquire data’ are essential to modelling and prediction in the natural and social sciences, engineering applications, and beyond. Optimal experimental design (OED) formalizes these questions and creates computational methods to answer them. This article presents a systematic survey of modern OED, from its foundations in classical design theory to current research involving OED for complex models. We begin by reviewing criteria used to formulate an OED problem and thus to encode the goal of performing an experiment. We emphasize the flexibility of the Bayesian and decision-theoretic approach, which encompasses information-based criteria that are well-suited to nonlinear and non-Gaussian statistical models. We then discuss methods for estimating or bounding the values of these design criteria; this endeavour can be quite challenging due to strong nonlinearities, high parameter dimension, large per-sample costs, or settings where the model is implicit. A complementary set of computational issues involves optimization methods used to find a design; we discuss such methods in the discrete (combinatorial) setting of observation selection and in settings where an exact design can be continuously parametrized. Finally we present emerging methods for sequential OED that build non-myopic design policies, rather than explicit designs; these methods naturally adapt to the outcomes of past experiments in proposing new experiments, while seeking coordination among all experiments to be performed. Throughout, we highlight important open questions and challenges.

97 MATHEMATICS AND COMPUTING↗

Seeing through noise in power laws

Despite widespread claims of power laws across the natural and social sciences, evidence in data is often equivocal. Modern data and statistical methods reject even classic power laws such as Pareto’s law of wealth and the Gutenberg–Richter law for earthquake magnitudes. We show that the maximum-likelihood estimators and Kolmogorov–Smirnov (K-S) statistics in widespread use are unexpectedly sensitive to ubiquitous errors in data such as measurement noise, quantization noise, heaping and censorship of small values. This sensitivity causes spurious rejection of power laws and biases parameter estimates even in arbitrarily large samples, which explains inconsistencies between theory and data. We show that logarithmic binning by powers of λ > 1 attenuates these errors in a manner analogous to noise averaging in normal statistics and that λ thereby tunes a trade-off between accuracy and precision in estimation. Binning also removes potentially misleading within-scale information while preserving information about the shape of a distribution over powers of λ, and we show that some amount of binning can improve sensitivity and specificity of K-S tests without any cost, while more extreme binning tunes a trade-off between sensitivity and specificity. We therefore advocate logarithmic binning as a simple essential step in power-law inference.

97 MATHEMATICS AND COMPUTING↗

COVID-19 dynamics across the US: A deep learning study of human mobility and social behavior

This paper presents a deep learning framework for epidemiology system identification from noisy and sparse observations with quantified uncertainty. The proposed approach employs an ensemble of deep neural networks to infer the time-dependent reproduction number of an infectious disease by formulating a tensor-based multi-step loss function that allows us to efficiently calibrate the model on multiple observed trajectories. The method is applied to a mobility and social behavior-based SEIR model of COVID-19 spread. The model is trained on Google and Unacast mobility data spanning a period of 66 days, and is able to yield accurate future forecasts of COVID-19 spread in 203 US counties within a time-window of 15 days. Interestingly, a sensitivity analysis that assesses the importance of different mobility and social behavior parameters reveals that attendance of close places, including workplaces, residential, and retail and recreational locations, has the largest impact on the effective reproduction number. Furthermore, the model enables us to rapidly probe and quantify the effects of government interventions, such as lock-down and re-opening strategies. Taken together, the proposed framework provides a robust workflow for data-driven epidemiology model discovery under uncertainty and produces probabilistic forecasts for the evolution of a pandemic that can judiciously provide information for policy and decision making. All codes and data accompanying this manuscript are available at https://github.com/PredictiveIntelligenceLab/DeepCOVID19.

60 APPLIED LIFE SCIENCES↗

Uncovering heterogeneous intercommunity disease transmission from neutral allele frequency time series

The COVID-19 pandemic has underscored the need for accurate epidemic forecasting to predict pathogen spread, evolution, and evaluate intervention strategies. Forecast reliability hinges on detailed knowledge of disease transmission across population segments, which may be inferred from contact surveys or mobility data. However, these indirect approaches make it difficult to estimate rare transmissions between socially or geographically distant communities. We show that the steep ramp-up of genome sequencing surveillance during the pandemic can be leveraged to directly identify transmission patterns between geographically defined communities. Our approach uses a hidden Markov model to infer the fraction of infections a community imports from others based on how rapidly allele frequencies in the focal community converge to those in the donor communities. Applying this method to SARS-CoV-2 sequencing data from England and the United States, we uncover networks of intercommunity transmission that reflect geographical relationships while exposing significant long-range interactions. The scaling of importation rate with distance is consistent across both countries, yet weaker than expected based on mobility data, highlighting limitations of indirect inference. We show that transmission patterns can change between waves of variants of concern and analyze how the inferred heterogeneity in intercommunity transmission impacts evolutionary forecasts. While applied here to geographically defined communities, our approach could be applied to those defined by other traits (e.g., age, socioeconomic status), provided time-series data can be stratified accordingly. Overall, our study highlights population genomic time series data as a crucial record of epidemiological interactions, which can be deciphered using tree-free inference methods.

Okada, Takashi [Department of Physics; University ↗

Cric searchable image database as a public platform for conventional pap smear cytology data

Amidst the current health crisis and social distancing, telemedicine has become an important part of mainstream of healthcare, and building and deploying computational tools to support screening more efficiently is an increasing medical priority. The early identification of cervical cancer precursor lesions by Pap smear test can identify candidates for subsequent treatment. However, one of the main challenges is the accuracy of the conventional method, often subject to high rates of false negative. While machine learning has been highlighted to reduce the limitations of the test, the absence of high-quality curated datasets has prevented strategies development to improve cervical cancer screening. The Center for Recognition and Inspection of Cells (CRIC) platform enables the creation of CRIC Cervix collection, currently with 400 images (1,376 × 1,020 pixels) curated from conventional Pap smears, with manual classification of 11,534 cells. This collection has the potential to advance current efforts in training and testing machine learning algorithms for the automation of tasks as part of the cytopathological analysis in the routine work of laboratories.

59 BASIC BIOLOGICAL SCIENCES↗

Seeing through noise in power laws

Despite widespread claims of power laws across the natural and social sciences, evidence in data is often equivocal. Modern data and statistical methods reject even classic power laws such as Pareto’s law of wealth and the Gutenberg–Richter law for earthquake magnitudes. We show that the maximum-likelihood estimators and Kolmogorov–Smirnov (K-S) statistics in widespread use are unexpectedly sensitive to ubiquitous errors in data such as measurement noise, quantization noise, heaping and censorship of small values. This sensitivity causes spurious rejection of power laws and biases parameter estimates even in arbitrarily large samples, which explains inconsistencies between theory and data. We show that logarithmic binning by powers of λ > 1 attenuates these errors in a manner analogous to noise averaging in normal statistics and that λ thereby tunes a trade-off between accuracy and precision in estimation. Binning also removes potentially misleading within-scale information while preserving information about the shape of a distribution over powers of λ, and we show that some amount of binning can improve sensitivity and specificity of K-S tests without any cost, while more extreme binning tunes a trade-off between sensitivity and specificity. We therefore advocate logarithmic binning as a simple essential step in power-law inference.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Modeling protected species distributions and habitats to inform siting and management of pioneering ocean industries: A case study for Gulf of Mexico aquaculture

Marine Spatial Planning (MSP) provides a process that uses spatial data and models to evaluate environmental, social, economic, cultural, and management trade-offs when siting (i.e., strategically locating) ocean industries. Aquaculture is the fastest-growing food sector in the world. The United States (U.S.) has substantial opportunity for offshore aquaculture development given the size of its exclusive economic zone, habitat diversity, and variety of candidate species for cultivation. However, promising aquaculture areas overlap many protected species habitats. Aquaculture siting surveys, construction, operations, and decommissioning can alter protected species habitat and behavior. Additionally, aquaculture-associated vessel activity, underwater noise, and physical interactions between protected species and farms can increase the risk of injury and mortality. In 2020, the U.S. Gulf of Mexico was identified as one of the first regions to be evaluated for offshore aquaculture opportunities as directed by a Presidential Executive Order. We developed a transparent and repeatable method to identify aquaculture opportunity areas (AOAs) with the least conflict with protected species. First, we developed a generalized scoring approach for protected species that captures their vulnerability to adverse effects from anthropogenic activities using conservation status and demographic information. Next, we applied this approach to data layers for eight species listed under the Endangered Species Act, including five species of sea turtles, Rice’s whale, smalltooth sawfish, and giant manta ray. Next, we evaluated four methods for mathematically combining scores (i.e., Arithmetic mean, Geometric mean, Product, Lowest Scoring layer) to generate a combined protected species data layer. The Product approach provided the most logical ordering of, and the greatest contrast in, site suitability scores. Finally, we integrated the combined protected species data layer into a multi-criteria decision-making modeling framework for MSP. This process identified AOAs with reduced potential for protected species conflict. These modeling methods are transferable to other regions, to other sensitive or protected species, and for spatial planning for other ocean-uses.

54 ENVIRONMENTAL SCIENCES↗

Identifying COVID-19 cases and extracting patient reported symptoms from Reddit using natural language processing

We used social media data from “covid19positive” subreddit, from 03/2020 to 03/2022 to identify COVID-19 cases and extract their reported symptoms automatically using natural language processing (NLP). We trained a Bidirectional Encoder Representations from Transformers classification model with chunking to identify COVID-19 cases; also, we developed a novel QuadArm model, which incorporates Question-answering, dual-corpus expansion, Adaptive rotation clustering, and mapping, to extract symptoms. Our classification model achieved a 91.2% accuracy for the early period (03/2020-05/2020) and was applied to the Delta (07/2021–09/2021) and Omicron (12/2021–03/2022) periods for case identification. We identified 310, 8794, and 12,094 COVID-positive authors in the three periods, respectively. The top five common symptoms extracted in the early period were coughing (57%), fever (55%), loss of sense of smell (41%), headache (40%), and sore throat (40%). During the Delta period, these symptoms remained as the top five symptoms with percent authors reporting symptoms reduced to half or fewer than the early period. During the Omicron period, loss of sense of smell was reported less while sore throat was reported more. Our study demonstrated that NLP can be used to identify COVID-19 cases accurately and extracted symptoms efficiently.

60 APPLIED LIFE SCIENCES↗

Scalable edge clustering of dynamic graphs via weighted line graphs

Timestamped relational datasets consisting of records (or connections) between pairs of entities are ubiquitous in network science. For applications like peer-to-peer communication, email, various social network interactions, and computer network security, it is useful to organize these records into groups based on how and when they are occurring. Weighted line graphs offer a natural way to model how records are related in such datasets but for large real-world graph topologies, building and utilizing the line graph is prohibitively expensive. Here, we present the framework to cluster the edges of a dynamic graph via the associated line graph that contains two major contributions. The first is a method to work with the line graph implicitly and the second is a distributed scale implementation of an agglomerative hierarchical graph clustering algorithm. We outline a novel hierarchical dynamic graph edge clustering approach that efficiently breaks massive relational datasets into small sets of edges containing events at various timescales. This is in stark contrast to traditional graph clustering algorithms that prioritize highly connected (clique-like) community structures. Our approach relies on constructing a sufficient subgraph of a weighted line graph and applying a hierarchical agglomerative clustering. This approach is related to scalable techniques from spatial clustering, nonlinear-dimension reduction, topological data analysis, and draws particular inspiration from HDBSCAN. As an edge clustering, this method yields an overlapping node clustering. Our algorithm is parallelizable and we demonstrate efficient clustering of a billion-scale, real-world dynamic graph into small edge sets that correlate in topology and time. The entire clustering process for a graph with tens of billions of edges takes just a few minutes of run time on 256 nodes of a distributed compute environment. We argue how the output of the edge clustering is useful for a multitude of data visualization and powerful machine learning tasks, both involving the original massive dynamic graph data and metadata associated with the nodes and edges. Finally, we describe how this approach can be extended to dynamic hypergraphs and dynamic graphs/hypergraphs with unstructured data living on vertices and edges.

Data Analysis↗