Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “statistical feature extraction”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Unstructured clinical notes within the 24 hours since admission predict short, mid & long-term mortality in adult ICU patients

Mortality prediction for intensive care unit (ICU) patients is crucial for improving outcomes and efficient utilization of resources. Accessibility of electronic health records (EHR) has enabled data-driven predictive modeling using machine learning. However, very few studies rely solely on unstructured clinical notes from the EHR for mortality prediction. In this work, we propose a framework to predict short, mid, and long-term mortality in adult ICU patients using unstructured clinical notes from the MIMIC III database, natural language processing (NLP), and machine learning (ML) models. Depending on the statistical description of the patients’ length of stay, we define the short-term as 48-hour and 4-day period, the mid-term as 7-day and 10-day period, and the long-term as 15-day and 30-day period after admission. We found that by only using clinical notes within the 24 hours of admission, our framework can achieve a high area under the receiver operating characteristics (AU-ROC) score for short, mid and long-term mortality prediction tasks. The test AU-ROC scores are 0.87, 0.83, 0.83, 0.82, 0.82, and 0.82 for 48-hour, 4-day, 7-day, 10-day, 15-day, and 30-day period mortality prediction, respectively. We also provide a comparative study among three types of feature extraction techniques from NLP: frequency-based technique, fixed embedding-based technique, and dynamic embedding-based technique. Lastly, we provide an interpretation of the NLP-based predictive models using feature-importance scores.

60 APPLIED LIFE SCIENCES↗

Domain Adaptive Graph Neural Networks for Constraining Cosmological Parameters Across Multiple Data Sets

Deep learning models have been shown to outperform methods that rely on summary statistics, like the power spectrum, in extracting information from complex cosmological data sets. However, due to differences in the subgrid physics implementation and numerical approximations across different simulation suites, models trained on data from one cosmological simulation show a drop in performance when tested on another. Similarly, models trained on any of the simulations would also likely experience a drop in performance when applied to observational data. Training on data from two different suites of the CAMELS hydrodynamic cosmological simulations, we examine the generalization capabilities of Domain Adaptive Graph Neural Networks (DA-GNNs). By utilizing GNNs, we capitalize on their capacity to capture structured scale-free cosmological information from galaxy distributions. Moreover, by including unsupervised domain adaptation via Maximum Mean Discrepancy (MMD), we enable our models to extract domain-invariant features. We demonstrate that DA-GNN achieves higher accuracy and robustness on cross-dataset tasks. Using data visualizations, we show the effects of domain adaptation on proper latent space data alignment. This shows that DA-GNNs are a promising method for extracting domain-independent cosmological information, a vital step toward robust deep learning for real cosmic survey data.

79 ASTRONOMY AND ASTROPHYSICS↗

Bridging the Gap between Cosmological Simulations with Graph Neural Networks and Domain Adaptation

Deep learning models have been shown to outperform methods that rely on summary statistics, like the power spectrum, in extracting information from complex cosmological data sets. However, due to differences in the subgrid physics implementation and numerical approximations across different simulation suites, models trained on data from one cosmological simulation show a drop in performance when tested on another. Similarly, models trained on any of the simulations would also likely experience a drop in performance when applied to observational data. Training on data from two different suites of the CAMELS hydrodynamic cosmological simulations, we examine the generalization capabilities of Domain Adaptive Graph Neural Networks (DA-GNNs). By utilizing GNNs, we capitalize on their capacity to capture structured scale-free cosmological information from galaxy distributions. Moreover, by including unsupervised domain adaptation via Maximum Mean Discrepancy (MMD), we enable our models to extract domain-invariant features. We demonstrate that DA-GNN achieves higher accuracy and robustness on cross dataset tasks (up to 28% better relative error and up to almost an order of magnitude better χ 2 ). Using data visualizations, we show the effects of domain adaptation on proper latent space data alignment. This shows that DA-GNNs are a promising method for extracting domain-independent cosmological information, a vital step toward robust deep learning for real cosmic survey data.

97 MATHEMATICS AND COMPUTING↗

Domain Adaptive Graph Neural Networks for Constraining Cosmological Parameters Across Multiple Data Sets

Deep learning models have been shown to outperform methods that rely on summary statistics, like the power spectrum, in extracting information from complex cosmological data sets. However, due to differences in the subgrid physics implementation and numerical approximations across different simulation suites, models trained on data from one cosmological simulation show a drop in performance when tested on another. Similarly, models trained on any of the simulations would also likely experience a drop in performance when applied to observational data. Training on data from two different suites of the CAMELS hydrodynamic cosmological simulations, we examine the generalization capabilities of Domain Adaptive Graph Neural Networks (DA-GNNs). By utilizing GNNs, we capitalize on their capacity to capture structured scale-free cosmological information from galaxy distributions. Moreover, by including unsupervised domain adaptation via Maximum Mean Discrepancy (MMD), we enable our models to extract domain-invariant features. We demonstrate that DA-GNN achieves higher accuracy and robustness on cross-dataset tasks (up to $28\%$ better relative error and up to almost an order of magnitude better $\chi^2$). Using data visualizations, we show the effects of domain adaptation on proper latent space data alignment. This shows that DA-GNNs are a promising method for extracting domain-independent cosmological information, a vital step toward robust deep learning for real cosmic survey data.

79 ASTRONOMY AND ASTROPHYSICS↗

Thresholding Analysis and Feature Extraction from 3D Ground Penetrating Radar Data for Noninvasive Assessment of Peanut Yield

This study explores the efficacy of utilizing a novel ground penetrating radar (GPR) acquisition platform and data analysis methods to quantify peanut yield for breeding selection, agronomic research, and producer management and harvest applications. Sixty plots comprising different peanut market types were scanned with a multichannel, air-launched GPR antenna. Image thresholding analysis was performed on 3D GPR data from four of the channels to extract features that were correlated to peanut yield with the objective of developing a noninvasive high-throughput peanut phenotyping and yield-monitoring methodology. Plot-level GPR data were summarized using mean, standard deviation, sum, and the number of nonzero values (counts) below or above different percentile threshold values. Best results were obtained for data below the percentile threshold for mean, standard deviation and sum. Data both below and above the percentile threshold generated good correlations for count. Correlating individual GPR features to yield generated correlations of up to 39% explained variability, while combining GPR features in multiple linear regression models generated up to 51% explained variability. The correlations increased when regression models were developed separately for each peanut type. This research demonstrates that a systematic search of thresholding range, analysis window size, and data summary statistics is necessary for successful application of this type of analysis. The results also establish that thresholding analysis of GPR data is an appropriate methodology for noninvasive assessment of peanut yield, which could be further developed for high-throughput phenotyping and yield-monitoring, adding a new sensor and new capabilities to the growing set of digital agriculture technologies.

54 ENVIRONMENTAL SCIENCES↗

An adaptive adversarial domain adaptation approach for corn yield prediction

Recently, statistical machine learning and deep learning methods have been widely explored for corn yield prediction. Though successful, machine learning models generated within a specific spatial domain often lose their validity when directly applied to new regions. To address this issue, we designed an unsupervised adaptive domain adversarial neural network (ADANN). Specifically, through domain adversarial training, the ADANN model reduced the impact of domain shift by projecting data from different domains into the same subspace. Also, the ADANN model was designed to be trained in an adaptive way, which guaranteed the model can learn the domain-invariant features and perform accurate yield prediction simultaneously. Informative variables including time-series vegetation indices and sequential weather observations were first collected from multiple data sources and aggregated to the county level. Then, we trained the ADANN model with the extracted features and corresponding reported county-level corn yield from the U.S. Department of Agriculture (USDA). Finally, the trained model was evaluated in four testing years 2016–2019. The U.S. corn belt was used as the study area and counties under study were grouped into two diverse ecological regions. Overall, the experimental results showed that the developed ADANN model had better performance than three other state-of-the-art machine learning models in both local experiments (train and test in the same region) and transfer experiments (train and test in different regions). As the first study using adversarial learning for crop yield prediction, this research demonstrates a novel solution for improving model transferability on crop yield prediction.

59 BASIC BIOLOGICAL SCIENCES↗

Local primordial non-Gaussian bias from time evolution

Primordial non-Gaussianity (PNG) is a signature of fundamental physics in the early Universe that is probed by cosmological observations. Here, it is well known that the local type of PNG generates a strong signal in the two-point function of large-scale structure tracers, such as galaxies. This signal, often termed “scale-dependent bias” is a generic feature of modulation of gravitational structure formation by a large-scale mode. It is less well appreciated that the coefficient controlling this signal, b ϕ , is closely connected to the time evolution of the tracer number density. This correspondence between time evolution and local PNG can be simply explained for a universal tracer whose mass function only depends on peak height and, more generally, for nonuniversal tracers in the separate universe picture, which we validate in simulations. We also describe how to recover the bias of tracers subject to a survey selection function and perform a simple demonstration on simulated galaxies. Since the local PNG amplitude in n-point statistics ($f$ NL ) is largely degenerate with the coefficient b ϕ , this proof of concept study demonstrates that Galaxy survey data can allow for more optimal and robust extraction of local PNG information from upcoming surveys.

Sullivan, James M. [University of California, Berk↗

Use of high-dimensional spectral data to evaluate organic matter, reflectance relationships in soils

Recent breakthroughs in remote sensing technology have led to the development of a spaceborne high spectral resolution imaging sensor, HIRIS, to be launched in the mid-1990s for observation of earth surface features. The effects of organic carbon content on soil reflectance over the spectral range of HIRIS, and to examine the contributions of humic and fulvic acid fractions to soil reflectance was evaluated. Organic matter from four Indiana agricultural soils was extracted, fractionated, and purified, and six individual components of each soil were isolated and prepared for spectral analysis. The four soils, ranging in organic carbon content from 0.99 percent, represented various combinations of genetic parameters such as parent material, age, drainage, and native vegetation. An experimental procedure was developed to measure reflectance of very small soil and organic component samples in the laboratory, simulating the spectral coverage and resolution of the HIRIS sensor. Reflectance in 210 narrow (10 nm) bands was measured using the CARY 17D spectrophotometer over the 400 to 2500 nm wavelength range. Reflectance data were analyzed statistically to determine the regions of the reflective spectrum which provided useful information about soil organic matter content and composition. Wavebands providing significant information about soil organic carbon content were located in all three major regions of the reflective spectrum: visible, near infrared, and middle infrared. The purified humic acid fractions of the four soils were separable in six bands in the 1600 to 2400 nm range, suggesting that longwave middle infrared reflectance may be useful as a non-destructive laboratory technique for humic acid characterization.

Henderson, T. L.↗

Training custom light curve models of SN Ia subpopulations selected according to host galaxy properties

ABSTRACT Type Ia supernova (SN Ia) cosmology analyses include a luminosity step function in their distance standardization process to account for an observed yet unexplained difference in the post-standardization luminosities of SNe Ia originating from different host galaxy populations [e.g. high-mass ($M \gtrsim 10^{10} \, {\rm M}_{\odot }$) versus low-mass galaxies]. We present a novel method for including host-mass correlations in the SALT3 (Spectral Adaptive Light curve Template 3) light curve model used for standardizing SN Ia distances. We split the SALT3 training sample according to host-mass, training independent models for the low- and high-host-mass samples. Our models indicate that there are different average Si ii spectral feature strengths between the two populations, and that the average spectral energy distribution of SNe from low-mass galaxies is bluer than the high-mass counterpart. We then use our trained models to perform an SN cosmology analysis on the 3-yr spectroscopically confirmed Dark Energy Survey SN sample, treating SNe from low- and high-mass host galaxies as separate populations throughout. We find that our mass-split models reduce the Hubble residual scatter in the sample, albeit at a low statistical significance. We do find a reduction in the mass-correlated luminosity step but conclude that this arises from the model-dependent re-definition of the fiducial SN absolute magnitude rather than the models themselves. Our results stress the importance of adopting a standard definition of the SN parameters (x0, x1, c) in order to extract the most value out of the light curve modelling tools that are currently available and to correctly interpret results that are fit with different models.

Taylor, G. (ORCID:0000000157563259)↗

Application of unsupervised deep learning to image segmentation and in-situ contact angle measurements in a CO 2 -water-rock system

Rock surface wettability is a critical property that regulates multiphase flows in porous media, which can be quantified using the surface contact angle (CA). X-ray micro-computed tomography (μCT) provides an effective approach to in-situ measurements of surface CAs. However, the CA measurement accuracy depends significantly on the quality of CT image segmentation, which is the clustering of CT pixels into separate phases. Inspired by this, we developed a deep learning (DL)-based CA measurement workflow. Motivated by the recent tremendous progress in unsupervised learning techniques and aiming to avoid expensive manual data annotations, an unsupervised DL pipeline for CT image segmentation was proposed and implemented, which includes unsupervised model training and post-processing. The unsupervised model training was driven by a novel loss function constrained with feature similarity and spatial continuity and implemented by iterative forward and backward paths; the former clustered the pixel-wise feature vectors extracted by convolution neural networks, whereas the latter updated the parameters using gradient descent. An over-segmentation strategy was adopted for model training. The post-processing steps based on agglomerative hierarchical clustering (AHC) were implemented to further merge the over-segmented model output to the desired cluster number, which is intended to improve the efficiency of image segmentation. The developed unsupervised DL pipeline was compared with other commonly-used image segmentation methods using pixel-wise and physics-based evaluation metrics on a synthetic raw-image dataset, which had a known ground truth. The unsupervised DL pipeline showed the best performance. Next, the segmented images were input to an automatic CA measurement tool, and the results were validated by comparisons with manual measurements. The CA values from the manual and automatic measurements showed similar distributions and statistical properties. The automatic measurement demonstrated a wider spectrum because of the much larger number of measurement data points. The primary novelty of the unsupervised DL pipeline developed in this study lies in the novel loss function and the over-segmentation strategy associated with AHC post-processing. Finally, the workflow has been proven an efficient tool for pore-scale wettability characterization, which has a wide range of applications in fundamental studies of multiphase flows in natural porous media, which have critical implications to geological carbon sequestration, hydrocarbon energy recovery, and contaminant transport in groundwater.

42 ENGINEERING↗

Climatological Processing of Radar Data for the TRMM Ground Validation Program

The Tropical Rainfall Measuring Mission (TRMM) satellite was successfully launched in November, 1997. The main purpose of TRMM is to sample tropical rainfall using the first active spaceborne precipitation radar. To validate TRMM satellite observations, a comprehensive Ground Validation (GV) Program has been implemented. The primary goal of TRMM GV is to provide basic validation of satellite-derived precipitation measurements over monthly climatologies for the following primary sites: Melbourne, FL; Houston, TX; Darwin, Australia; and Kwajalein Atoll, RMI. As part of the TRMM GV effort, research analysts at NASA Goddard Space Flight Center (GSFC) generate standardized TRMM GV products using quality-controlled ground-based radar data from the four primary GV sites as input. This presentation will provide an overview of the TRMM GV climatological processing system. A description of the data flow between the primary GV sites, NASA GSFC, and the TRMM Science and Data Information System (TSDIS) will be presented. The radar quality control algorithm, which features eight adjustable height and reflectivity parameters, and its effect on monthly rainfall maps will be described. The methodology used to create monthly, gauge-adjusted rainfall products for each primary site will also be summarized. The standardized monthly rainfall products are developed in discrete, modular steps with distinct intermediate products. These developmental steps include: (1) extracting radar data over the locations of rain gauges, (2) merging rain gauge and radar data in time and space with user-defined options, (3) automated quality control of radar and gauge merged data by tracking accumulations from each instrument, and (4) deriving Z-R relationships from the quality-controlled merged data over monthly time scales. A summary of recently reprocessed official GV rainfall products available for TRMM science users will be presented. Updated basic standardized product results and trends involving monthly accumulation, Z-R relationship, and gauge statistics for each primary GV site will be also displayed.

Kulie, Mark↗

Algorithmic Classification of Raman Spectra Biosignatures: Improving Life Detection Confidence

“Agnostic” biosignatures – indicators of life (or the absence of life), independent of a particular biochemistry – are increasingly considered a high standard for life detection. The Ladder of Life Detection (2018) called for investigating how combinations of independent and different potential biosignatures affect confidence. To address this gap, statistical classification of elemental abundances, isotopic fractionation, and reflectance spectroscopy (VNIR) has been implemented. Raman spectroscopy, highly desirable due to its wide availability, has the potential to improve this predictive power. This work implemented biosignature classification algorithms on Raman data alone, in preparation for combination with the other data types. Raman spectroscopy data was collected from published databases and papers as part of a manually curated dataset of “indicative” and “non-indicative of life” samples. These currently include 61 non-indicative samples (meteorites, magnetite); 3 indicative living samples (bacteria); 20 indicative non-living samples (chalk, bone); and 12 indicative mixed (with non-indicative material) samples (soil, microbial mats). Laboratory work is ongoing to characterize additional samples, particularly a greater breadth of mixed systems. Spectra were interpolated, filtered with the Savitzsky-Golay filter, and de-noised. For a preliminary examination, agnostic features were manually extracted including mean intensity, number of peaks, and mean peak width. Different peak prominences and filtering polynomials were used to refine features. Classification algorithms were implemented: k-nearest neighbors (KNN), logistic regression (LR), linear support vector machines (SVM), random forest (RF), Gaussian naïve bayes (GNB). Lastly, Monte Carlo simulations on 1,000 50%-train-test-splits were used to validate classification performance and feature significance. The preliminary feature set achieved its highest AUC of 0.52 with LR, with no strongly discriminatory features. Work to improve feature extraction, such as through deep learning with back propagation, is planned. In future work, the Raman data will be combined with the other data types, and potentially new data types such as enantiomeric excess. This project was partially supported through the NASA Ames Project EXcellence (APEX) incubator program.

Astrobiology↗

Estimations of ABL fluxes and other turbulence parameters from Doppler lidar data

Techniques for extraction boundary layer parameters from measurements of a short-pulse CO2 Doppler lidar are described. The measurements are those collected during the First International Satellites Land Surface Climatology Project (ISLSCP) Field Experiment (FIFE). By continuously operating the lidar for about an hour, stable statistics of the radial velocities can be extracted. Assuming that the turbulence is horizontally homogeneous, the mean wind, its standard deviations, and the momentum fluxes were estimated. Spectral analysis of the radial velocities is also performed from which, by examining the amplitude of the power spectrum at the inertial range, the kinetic energy dissipation was deduced. Finally, using the statistical form of the Navier-Stokes equations, the surface heat flux is derived as the residual balance between the vertical gradient of the third moment of the vertical velocity and the kinetic energy dissipation. Combining many measurements would normally reduce the error provided that, it is unbiased and uncorrelated. The nature of some of the algorithms however, is such that, biased and correlated errors may be generated even though the raw measurements are not. Data processing procedures were developed that eliminate bias and minimize error correlation. Once bias and error correlations are accounted for, the large sample size is shown to reduce the errors substantially. The principal features of the derived turbulence statistics for two case studied are presented.

Gal-Chen, Tzvi↗

FY21 Progress Report: SRNL Analysis of ICCWR LCM and WAMS data for Corrosion and Cracking

The development of algorithms for machine learning and data analysis for the 3013 Surveillance Program is a collaborative effort by the Savannah River National Laboratory (SRNL) and the University of South Carolina (USC). For corrosion detection, Laser Confocal Microscope (LCM) or Wide Area 3D Measurement System (WAMS) data is extracted from large binary files, with software written to convert the data to physical attributes (e.g., height, color and grayscale values; all as functions of a location in a plane projection). A user-friendly Matlab Graphical User Interface (GUI) that reads data from either LCM or WAMS files was developed to integrate input data with software developed for processing and evaluation. The GUI can selectively download binary data, interrogate data attributes, label data, flag significant features, execute Machine Learning (ML) algorithms, output parameters for trained ML algorithms, report ML model accuracy with respect to labeled data, and generate graphical representations for various analyses. Features can be called out by user-specified thresholds, manual labeling or machine learning algorithms when they have been completed. The ability to rapidly label data is important because of the volume of data required for training machine learning algorithms. The GUI has the flexibility to allow addition of improved ML algorithms, methods for data visualization, and statistical computations. Statistical analyses via the GUI include areas of pits within a defined range of pit depths, correlations between Red-Green-Blue (RGB) or grayscale intensity and relative surface height, covariances between values associated with features, and feature histograms. The development of supervised machine learning algorithms, however, has been hindered by a lack of training data. The machine learning algorithms for crack identification are being refined but require improvements to the true positive rate for crack detection. This shortcoming is an artifact of the limited training data currently available, perhaps more so than the structure of the neural networks. At present, the best results are had from a consensus over an ensemble of randomly generated Deep Neural Network (DNN) or Convolutional Neural Network (CNN) algorithms. Although the consensus accuracy method has yielded optimum true positive and true negative rates in excess of 80%, additional validation testing is necessary. In addition to the suite of LCM data that was initially used, and which represents the majority of the work presented in this report, WAMS image data was also reviewed at a preliminary level. The review included a comparison between image resolution and dynamic range for each method. WAMS (ZON file) image data was found to have a pixel pitch of 3.69μm compared to 1 μm for the LCM (vk4 file) data, which implies a lower resolution for the WAMS images. Conversely, the ratio of dynamic range of the WAMS data to the LCM data was approximately 41:20 for height data, suggesting that information from WAMS should more accurately determine the depth of pits. At present, the significance of the greater dynamic range of the WAMS data relative to the LCM data has not yet been evaluated.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

Intelligent System Development Using a Rough Sets Methodology

The purpose of this research was to examine the potential of the rough sets technique for developing intelligent models of complex systems from limited information. Rough sets a simple but promising technology to extract easily understood rules from data. The rough set methodology has been shown to perform well when used with a large set of exemplars, but its performance with sparse data sets is less certain. The difficulty is that rules will be developed based on just a few examples, each of which might have a large amount of noise associated with them. The question then becomes, what is the probability of a useful rule being developed from such limited information? One nice feature of rough sets is that in unusual situations, the technique can give an answer of 'I don't know'. That is, if a case arises that is different from the cases the rough set rules were developed on, the methodology can recognize this and alert human operators of it. It can also be trained to do this when the desired action is unknown because conflicting examples apply to the same set of inputs. This summer's project was to look at combining rough set theory with statistical theory to develop confidence limits in rules developed by rough sets. Often it is important not to make a certain type of mistake (e.g., false positives or false negatives), so the rules must be biased toward preventing a catastrophic error, rather than giving the most likely course of action. A method to determine the best course of action in the light of such constraints was examined. The resulting technique was tested with files containing electrical power line 'signatures' from the space shuttle and with decompression sickness data.

Anderson, Gray T.↗

Dust and the intrinsic spectral index of quasar variations: hints of finite stress at the innermost stable circular orbit

ABSTRACT We present a study of 9 242 spectroscopically confirmed quasars with multiepoch ugriz photometry from the SDSS Southern Survey. By fitting a separable linear model to each quasar’s spectral variations, we decompose their five-band spectral energy distributions into variable (disc) and non-variable (host galaxy) components. In modelling the disc spectra, we include attenuation by dust on the line of sight through the host galaxy to its nucleus. We consider five commonly used attenuation laws, and find that the best description is by dust similar to that of the Small Magellanic Cloud, inferring a lack of carbonaceous grains from the relatively weak 2175-Å absorption feature. We go on to construct a composite spectrum for the quasar variations spanning 700–8000 Å. By varying the assumed power-law Lν ∝ να spectral slope, we find a best-fitting value α = 0.71 ± 0.02, excluding at high confidence the canonical Lν ∝ ν1/3 prediction for a steady-state accretion disc with a T ∝ r−3/4 temperature profile. The bluer spectral index of the observed quasar variations instead supports the model of Agol & Krolik, and Mummery & Balbus, in which a steeper temperature profile, T ∝ r−7/8, develops as a result of finite magnetically induced stress at the innermost stable circular orbit extracting energy and angular momentum from the black hole spin.

79 ASTRONOMY AND ASTROPHYSICS↗

The Multiplatform Precipitation Feature (MPF) Database: Synthesizing Satellite and Ground-Based Precipitation and Lightning Datasets for Convective Studies

NASA’s Lightning Imaging Sensor (LIS) and the Global Precipitation Measurement (GPM) mission have contributed a wealth of data toward global lightning and precipitation studies, respectively. Combining lightning and precipitation datasets leverages their unique insights into deep convective processes that inform about characteristics of convection and its intensity. Recent efforts to synthesize the LIS and GPM datasets prepare the opportunity for unprecedented large-scale, value-added multiplatform analyses of convection. This data synthesis proof-of-concept study elaborates on the creation of a database of reflectivity-based multiplatform precipitation features (MPFs) that capture a combination of information extracted from spatiotemporally coincident lightning and precipitation data within individual storm features. The space-based GPM Dual-frequency Precipitation Radar (DPR) provides a record of precipitation data, while the GPM Validation Network (VN) additionally incorporates ground-based polarimetric Doppler radar data to provide microphysical and kinematic context to DPR data. The LIS instrument onboard the International Space Station has contributed lightning observations since 2017. MPFs encapsulating information from these datasets are created from isolated regions of filtered, smoothed DPR reflectivity data to which ellipses are fit. Each MPF includes feature location, size, and eccentricity information as well as summary reflectivity characteristics. They also include summaries of precipitation microphysics and derived three-dimensional wind available from ground-based radar data. LIS data provides standard lightning characteristics such as flash count and density to each MPF as well as other informative metrics such as flash area and radiance. Each MPF file includes information about the original data from which the MPF and its characteristics were determined, allowing end-user reconstruction of the ellipse and deeper “level I” analysis of captured data. This database of VN-LIS MPFs enables broad statistical analysis of the relationships between the microphysical, kinematic, and electrical properties of convection. Preliminary results from a demonstration of the database will be described as well as ongoing efforts and avenues for future work.

Lightning↗

Cosmic Complexity

What explains the extraordinary complexity of the observed universe, on all scales from quarks to the accelerating universe? My favorite explanation (which I certainty did not invent) ls that the fundamental laws of physics produce natural instability, energy flows, and chaos. Some call the result the Life Force, some note that the Earth is a living system itself (Gaia, a "tough bitch" according to Margulis), and some conclude that the observed complexity requires a supernatural explanation (of which we have many). But my dad was a statistician (of dairy cows) and he told me about cells and genes and evolution and chance when I was very small. So a scientist must look for me explanation of how nature's laws and statistics brought us into conscious existence. And how is that seemll"!gly Improbable events are actually happening a!1 the time? Well, the physicists have countless examples of natural instability, in which energy is released to power change from simplicity to complexity. One of the most common to see is that cooling water vapor below the freezing point produces snowflakes, no two alike, and all complex and beautiful. We see it often so we are not amazed. But physlc!sts have observed so many kinds of these changes from one structure to another (we call them phase transitions) that the Nobel Prize in 1992 could be awarded for understanding the mathematics of their common features. Now for a few examples of how the laws of nature produce the instabilities that lead to our own existence. First, the Big Bang (what an insufficient name!) apparently came from an instability, in which the "false vacuum" eventually decayed into the ordinary vacuum we have today, plus the most fundamental particles we know, the quarks and leptons. So the universe as a whole started with an instability. Then, a great expansion and cooling happened, and the loose quarks, finding themselves unstable too, bound themselves together into today's less elementary particles like protons and neutrons, liberating a little energy and creating complexity. Then, the expanding universe cooled some more, and neutrons and protons, no longer kept apart by immense temperatures, found themselves unstable and formed helium nuclei. Then, a little more cooling, and atomic nuclei and electrons were no longer kept apart, and the universe became transparent. Then a little more cooling, and the next instability began: gravitation pulled matter together across cosmic distances to form stars and galaxies. This instability is described as a "negative heat capadty" in which extracting energy from a gravitating system makes it hotter -- clearly the 2nd law of thermodynamics does not apply here! (This is the physicist's part of the answer to e e cummings' question: what is the wonder that's keeping the stars apart?) Then, the next instability is that hydrogen and helium nuclei can fuse together to release energy and make stars burn for billions of years. And then at the end of the fuel source, stars become unstable and explode and liberate the chemical elements back into space. And because of that, on planets like Earth, sustained energy flows support the development of additional instabilities and all kinds of complex patterns. Gravitational instability pulls the densest materials into the core of the Earth, leaving a thin skin of water and air, and makes the interior churn incessantly as heat flows outwards. And the heat from the sun, received mostly near the equator and flowing towards the poles, supports the complex atmospheric and oceanic circulations. And because or that, the physical Earth is full of natural chemical laboratories, concentrating elements here, mixing them there, raising and lowering temperatures, ceaselessly experimenting with uncountable events where new instabilities can arise. At least one of them was the new experiment called life. Now that we know that there are at least as many planets as there are stars, it is hard to imagine that nature's ceasess experimentation would not be able to produce life elsewhere -- but we don't know for sure. And life went on to cause new Instabilities, constantly evolving, with living things in an extraordinary range of environments, changing the global environment, with boom-and-bust cycles. with predators for every kInd of prey, with criminals for every possible crime, with governments to prevent them, and instabilities of the governments themselves. One of the instabilities Is that humans demand new weapons and new products of all sort, leading to serious investments in science and technology. So the natural/human world of competition and combat is structured to lead to advanced weaponry and cell phones. So here we are In 2012, with people writing essays and wondering whether their descendents will be artificial life forms travelling back into space. And, pondering what are the origins of those forces of nature that give rise to everything. Verllnde has argued that gravitation, the one force that has so far resisted our efforts at a Quantum description, is not even a fundamental force, but is itself it a statistical force, like osmosis. What an amazing turn of events! But after all I've just said, I should not be surprised a bit.

Mather, John C.↗