Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “social networking”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Evolving efforts to maintain and improve XPS analysis quality in an era of increasingly diverse uses and users

Based on literature analysis, X-ray photoelectron spectroscopy (XPS) use continues to increase exponentially. This increased use is accompanied by anecdotal reports and systematic analyses indicating a growing presence of significantly flawed data analyses. Recognition of this problem within the surface analysis community has increased with an understanding that both inexperienced users and increased use of XPS outside the surface analysis community contribute to the problem. The XPS community has initiated several efforts to help address the problem, which is not unique to XPS. This paper describes some of the specific problems identified and some of the community efforts intended to address them. Here, we describe activities focused on three specific issues: (i) requests for detailed guides and protocols and bite-sized versions of information for non-experts, (ii) incomplete data and analysis reporting, and (iii) the high rate of peak fitting problems. A 2019 survey identified the need for guides, protocols, and standards to assist XPS users. One set of such guides has been published, and another is being assembled. Providing incremental bites of useful information is the goal of a series of papers on specific challenges to surface analysis with example solutions has been initiated as Notes and Insights papers in Surface and Interface Analysis. Examination of XPS-containing papers finds that information to establish the credibility and reproducibility of XPS results is often very incomplete. Unfortunately, ISO and ASTM standards require an amount of parameter reporting that seems excessive and unrealistic for many research publications. Initial approaches to develop and distribute a graded approach to parameter reporting are briefly described. Multiple efforts are underway to address the high rate of problems associated with photoelectron peak fitting. These include guides to peak fitting, guides to peak identification and fitting for specific elements, and the development of a peak fitting social network. The fitting social network is designed to facilitate interactions between new and experienced XPS users; analysts trying to fit XPS data (for publication or other reasons) can ask questions and establish dynamic conversations. Encouraging and enabling high-quality XPS analysis and reporting requires several different types of effort from all members of the surface and interface analysis community.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

A Data Processing Pipeline To Extract A Knowledge Graph From Heterogeneous Data For Socio-technical Analysis Of Critical Infrastructure Influence

The code is written in Python and consists of the following pipeline that is implemented in Apache Airflow. This pipeline intends to understand the companies that are directly or indirectly involved with a type of critical infrastructure system at some point in that system's lifecycle. The pipeline takes a configuration file that specifies a list of initial companies to consider, a geographic region of interest, and a set of SEC form types as well as other data sources (e.g. CrunchBase) from which to extract entities and relations. There are four main components to this pipeline as currently implemented: Entity Extraction, Network Construction, Analysis, and Visualization. First, Entity Extraction, is implemented as the `topear-extract_organizations` Apache Airflow workflow. Given an initial query that specifies a geographic region of interest and a time interval, the software will extract CI facilities of interest and organizations that have a direct influence relationship to those facilities (e.g. ownership). During the course of the LDRD, we focused on Electric Vehicle charging stations and this information is available via the Department of Energy (DOE) database on fueling stations maintained by NREL. Within the context of the DOE CESER project, we have focused on Battery Energy Storage Systems (BESS). Second, the Network Extraction component will iteratively construct a social network graph given the set of organizations and people extracted in the previous step. Organizations (and eventually People if desired) are then fed as a query to the `topgear-construct_social_network` Apache Airflow workflow which given a set of initial companies and data sets (e.g. SEC EDGAR form types, OpenCorporates, Crunchbase). This Airflow workflow will iteratively query such data sources to discover relationships with new organizations and people. For example, this module can iteratively query SEC EDGAR for metadata that documents the number of each type of form for the given set of companies and their location. This forms metadata represents a catalog of data sources from SEC EDGAR for the extracted social network knowledge graph. The pipeline then downloads these forms from the website and saves them in a build directory for further processing. These documents are then parsed for entities and relations. Again, we note that in additional to SEC data sources, this step can also pull in information on organizations via API services such as CrunchBase and OpenCorporates or bulk data sources. At the end of this step, the resultant social network, the Critical Infrastructure network, and the edges that encode relationships between organizations and CI facilities, form the Adversarial Socio-Technical Network (ASTN) that informs the analysis. Third, the Analysis component processes these generated ASTN. Previously, that has included the ability to compare prevalence of different vendors for a given infrastructure component type across different regions as well as identify common public and private investors across those vendors. This was demonstrated for EV Charging Stations across several different metropolitan areas within an IEEE PES GridEdge publication. More recently, we have looked at ways to identify infrastructure owners and operators of BESS with the most nameplate capacity across different states as well as other indictors of risk resulting from changes in ownership over time. Finally, the Visualization component consists of an HTML/CSS/JS framework by which users can interact geospatial, operational, and organizational relationships across a given portfolio of Critical Infrastructure facilities. The objective is to provide a library of UI/UX modules that can be repurposed for stakeholder-specific dashboards. All of the modules are related via a common event model that enables UI actions in one view to percolate across the other views.

Weaver, Gabriel [Idaho National Laboratory (INL), ↗

A Data Processing Pipeline To Extract A Knowledge Graph From Sec Documents For Socio-technical Analysis Of Critical Infrastructure Influence

The code is written in Python and consists of the following pipeline that is implemented in Apache Airflow. This pipeline intends to understand the companies that are directly or indirectly involved with a type of critical infrastructure system at some point in that system's lifecycle. The pipeline takes a configuration file that specifies a list of initial companies to consider, a geographic region of interest (disk) expressed as a latitude/longitude point and distance, and a set of SEC form types from which to extract entities and relations. There are three main components to this pipeline as currently implemented: Social Network Extraction, Critical Infrastructure Network Extraction, and Inference and Fusion. First, Social Network Extraction, implemented as the `organizations_sec` component of the workflow graph queries the SEC EDGAR webservice using the list of initial companies from the configuration file. Given this, it extracts metadata that documents the number of each type of form for the given set of companies and their location. This forms metadata represents a catalog of data sources for the extracted social network knowledge graph. The pipeline then downloads these forms from the website and saves them in a build directory for further processing. These documents are then parsed for entities and relations. Second, the Critical Network Extraction component extracts entities and relations for a critical infrastructure sector. Currently, we focus on Electric Vehicle charging stations and this information is available via the Department of Energy (DOE) database on fueling stations maintained by NREL. Third, the Inference and Fusion component relates the social network graph to the critical infrastructure graph in order to understand the impact of a company within a geographic region. Relations include ownership of the EV Charging Station asset as well as maintenance/ownership of the EV payment networks. The fused network can be represented in many ways and currently we emit a knowledge graph.

Weaver, GabrielA.↗

How deep to dig: effects of web-scraping search depth on hyperlink network analysis of environmental stewardship organizations

Abstract Social network analysis (SNA) tools and concepts are essential for addressing many environmental management and sustainability issues. One method to gather SNA data is to scrape them from environmental organizations’ websites. Web-based research can provide important opportunities to understand environmental governance and policy networks while potentially reducing costs and time when compared to traditional survey and interview methods. A key parameter is ‘search depth,’ i.e., how many connected pages within a website to search for information. Existing research uses a variety of depths and no best practices exist, undermining research quality and case study comparability. We therefore analyze how search depth affects SNA data collection among environmental organizations, if results vary when organizations have different objectives, and how search depth affects social network structure. We find that scraping to a depth of three captures the majority of relevant network data regardless of an organization’s focus. Stakeholder identification (i.e., who is in the network) may require less scraping, but this might under-represent network structure (i.e., who is connected). We also discuss how scraping web-pages of local programs of larger organizations may lead to uncertain results and how our work can combine with mixed methods approaches.

Sayles, Jesse S. (ORCID:0000000218378920)↗

Secure Peer-to-Peer Networks for Scientific Information Sharing

The most common means of remote scientific collaboration today includes the trio of e-mail for electronic communication, FTP for file sharing, and personalized Web sites for dissemination of papers and research results. With the growth of broadband Internet, there has been a desire to share large files (movies, files, scientific data files) over the Internet. Email has limits on the size of files that can be attached and transmitted. FTP is often used to share large files, but this requires the user to set up an FTP site for which it is hard to set group privileges, it is not straightforward for everyone, and the content is not searchable. Peer-to-peer technology (P2P), which has been overwhelmingly successful in popular content distribution, is the basis for development of a scientific collaboratory called Scientific Peer Network (SciPerNet). This technology combines social networking with P2P file sharing. SciPerNet will be a standalone application, written in Java and Swing, thus insuring portability to a number of different platforms. Some of the features include user authentication, search capability, seamless integration with a data center, the ability to create groups and social networks, and on-line chat. In contrast to P2P networks such as Gnutella, Bit Torrent, and others, SciPerNet incorporates three design elements that are critical to application of P2P for scientific purposes: User authentication, Data integrity validation, Reliable searching SciPerNet also provides a complementary solution to virtual observatories by enabling distributed collaboration and sharing of downloaded and/or processed data among scientists. This will, in turn, increase scientific returns from NASA missions. As such, SciPerNet can serve a two-fold purpose for NASA: a cost-savings software as well as a productivity tool for scientists working with data from NASA missions.

Karimabadi, Homa↗

Trust based attachment

In social systems subject to indirect reciprocity, a positive reputation is key for increasing one’s likelihood of future positive interactions. The flow of gossip can amplify the impact of a person’s actions on their reputation depending on how widely it spreads across the social network, which leads to a percolation problem. To quantify this notion, we calculate the expected number of individuals, the “audience”, who find out about a particular interaction. For a potential donor, a larger audience constitutes higher reputational stakes, and thus a higher incentive, to perform “good” actions in line with current social norms. For a receiver, a larger audience therefore increases the trust that the partner will be cooperative. This idea can be used for an algorithm that generates social networks, which we call trust based attachment (TBA). TBA produces graphs that share crucial quantitative properties with real-world networks, such as high clustering, small-world behavior, and powerlaw degree distributions. We also show that TBA can be approximated by simple friend-of-friend routines based on triadic closure, which are known to be highly effective at generating realistic social network structures. Therefore, our work provides a new justification for triadic closure in social contexts based on notions of trust, gossip, and social information spread. These factors are thus identified as potential significant influences on how humans form social ties.

59 BASIC BIOLOGICAL SCIENCES↗

Sun-Earth Day: Reaching the Education Audience by Informal Means

For ten years the Sun-Earth Day program has promoted Heliophysics education to ever larger audiences through events centered on attractive annual themes. What originally started out as a one day event quickly evolved into a series of programs and events that occur throughout the year culminating with a celebration on or near the Spring Equinox. The events are often formal broadcasts or webcasts seeking to convey the science behind the latest solar-terrestrial mission discoveries. This has been quite successful, but it is clear that the younger generation increasingly depends on social networking approaches and informal news transmission for learning what is happening in the world around them. For 2010, the Sun-Earth Day team put emphasis on using informal approaches to bring the theme to the audience. The main event, a webcast from the NASA booth at the National Science Teachers Association (NSTA) annual meeting by the NASA EDGE group, took a lighthearted and offbeat approach to interviewing scientists and educators about Heliophysics news. NASA EDGE programs are unscripted and unpredictable, and that represents a different approach to getting the message across. The webcast was supplemented by a number of social networking avenues. The Sun-Earth Day program explored a wide range of social media applications including Facebook, Twitter, NING, podcasting, iPhone apps, etc. Each of these offers unique and effective methods to promote Heliophysics content and mission related highlights. The facebook site was quite popular and message posting there told the Sun-Earth Day story piece by piece. The same could be said of twittering and the tweetup held at the NSTA site. Has all of this been effective? Results are still being gathered, but anecdotal responses from the world seem very positive. What other methods might be used in the future to bring the science to a personal hands-on, interactive experience? Outcomes: Participants will: (1) Be introduced to the Sun-Earth Day program and its evolution through a decade of programs; (2) Hear about the methods used to communicate and educate through the years and how well they have worked; and (3) Be acquainted with the latest usage of social networking and informal education approaches and how well they have worked

Thieman, J.↗

NSCOR for Evaluating Risk Factors and Biomarkers for Adaptation and Resilience to Spaceflight: Emotional Valence and Social Processes in ICC/ICE Environments

Space exploration class missions, such as a mission to Mars, will require optimization of human performance, adaptability, and resilience. This NASA Specialized Center of Research (NSCOR) utilizes the NIMH Research Domain Criteria (RDoC) framework to identify biological and behavioral markers of individual social adaptation and emotional resilience (as well as vulnerability) to spaceflight-relevant stressors such as living in extended isolation. The overarching goal of this NSCOR is to obtain novel information to help identify biomarkers of individuals who are resilient and/or adaptable to the stressors of isolated, confined, and controlled (ICC) and isolated, confined, and extreme (ICE) environments.A total of N=90 healthy adult astronaut surrogates are being studied in three spaceflight-analog environments: (1) n=40 healthy adults in the Isolation and Confinement Analog Research Unit (ICARUS), an ICC at the University of Pennsylvania, during 7-day missions, for a target total of 280 subject days; (2) n=32 healthy adult astronaut surrogates studied in NASA’s Human Exploration Research Analog (HERA), an ICC at Johnson Space Center during 45-day missions, for a target total of 2,112 subject days; and (3) n=18 healthy adults in the Alfred-Wegener-Institute’s Neumayer Station III, an ICE in Antarctica, during 14-month missions, for a target total of 7,560 subject days. Dr. Nindl’s Laboratory at the University of Pittsburgh is analyzing a priori selected protein biomarkers in blood, saliva, and urine. Complementary rodent models of exposure to early life stressors, confinement, and isolation are being evaluated at Dr. Hensch’s Laboratory to further validate the neurobehavioral and biological findings from the human studies.Given the inconsistency and varied definition of resilience/adaptation in the scientific literature, the NSCOR team developed a composite resilience/adaptation measure that reflects the most relevant outcomes to resilience/adaptation across psychosocial and neurobehavioral functions, as well as neurocognitive and spaceflight-relevant operational performance. To achieve this, group consensus was attained from subject matter experts to produce a rank-order of importance for each input variable. The final resilience/adaptation score included 36 variables that were collected across spaceflight analogs. As of 10/1/2021, the NSCOR project acquired data on n=27 subjects at ICARUS, n=16 at HERA, and n=18 at Neumayer. The COVID-19 pandemic delayed data acquisition at ICARUS and HERA.Among subjects studied to date, 99% of neural and neurobehavioral data (e.g., neuroimaging for structure and function, behavioral measures) as well as blood, saliva, and urine for biochemical assays have been acquired. For rodent models, Dr. Hensch’s laboratory has established biochemical and behavioral parameters reflecting confinement stress in social networks of mice for comparison to stress responses in the human spaceflight analog environments.Group social behaviors were measured with a Social Network Analysis (SNA) approach to define objective parameters associated with sociability and its plasticity by sex. This analytic approach may help identify a network of individuals who are more effective teammates or more likely to generate new social relationships. Data acquisition, biomarker assessment, and data quality control will continue through September 2022.

D F Dinges↗

Evaluating Direct and Indirect Influence on EV Charging Stations Across the US

The adoption of new technology for electric vehicles (EV) and mobility applications can bring underappreciated vulnerabilities to the power grid. One area of potential fraud and adversarial influence is through the business ecosystem of startups that own and deploy EV technology. Yet, there are no models or analyses that map the network of organizations and people that have direct and indirect influence over technologies currently deployed in the grid. To fill this gap, we develop a multilayer network model to measure direct and indirect influence on EV charging stations. First, we create and adversarial socio-technical network (ASTN) model via a data fusion pipeline for different US regions of interest (ROI). Then, we develop an integrated ASTN for Chicago, Los Angeles, New York, and Philadelphia. We rank EV charging companies direct influence within each geographic region as well as indirect influence via social network analysis. While some companies have strong direct and indirect influence (i.e., ChargePoint) others show a mismatch between their influence over charging stations and their position within the social network. For example, Tesla has strong direct influence on stations and weak indirect influence over competitors. In contrast, 7Charge has weak direct influence over stations, but strong indirect influence over competitors.

29 - ENERGY PLANNING, POLICY AND ECONOMY↗

Exploring the Landscape of Distributed Graph Clustering on Leadership Supercomputers

The rapid growth of large-scale datasets in fields like biology and social networks has driven the need for advanced graph analytics techniques. Community detection, a fundamental task in graph analytics, identifies closely connected groups of nodes within a network, providing valuable insights across various disciplines. This study focuses on two classic community detection methods, the Louvain algorithm and Markov Clustering (MCL), and evaluates the performance of two prominent distributed community detection algorithms: HiPDPL-GPU, our prior implementation, and HipMCL. We conduct experiments on GPU-accelerated heterogeneous HPC systems, Summit and Frontier, to assess their performance under varying conditions. Our objective is to identify the strengths and weaknesses of these algorithms in terms of scalability, and quality of solutions. We evaluate these algorithms on a diverse set of 70+ networks spanning 13 domains, with sizes ranging up to 4.2 billion edges. Our results demonstrate that HiPDPL-GPU consistently outperforms HipMCL, especially for large-scale networks. HiPDPL-GPU achieves significantly faster runtimes (47x to 1439x), higher modularity scores, and improved scalability. These findings highlight HiPDPL-GPU as a promising solution for efficient and effective large-scale graph analytics in diverse application domains, and provide insights into the feasibility of using MCL-based approaches for certain application domains.

Community detection, graph algorithms↗

HBMax: Optimizing Memory Efficiency for Parallel Influence Maximization on Multicore Architectures

The goal of influence maximization is to select k most-influential vertices or seeds in a network, where influence is defined by a given diffusion process. The problem has a number of important applications such as viral marketing, information spread, and epidemic control. Although computing optimal seed set is NP-Hard, due to the submodular nature of the problem efficient approximation algorithms exist. However, even state-of-the-art parallel implementations are limited by a sampling step that incurs large memory footprints. This in turn limits the problem size reach and approximation quality. In this work, we study the memory footprint of the sampling process collecting reverse reachability information in the IMM algorithm over large real-world social networks. We present an adaptive and memory-efficient optimization approach for a state-of-the-art multi-threaded parallel influence maximization algorithm. Our approach,HuffMax, uses a portion of the reverse reachable (RR) sets collected by the algorithm to learn the characteristics of the graph. Then, it compresses the intermediate reverse reachability information with Huffman coding, and queries directly on the compressed data to preserve the memory savings obtained through compression. We also propose an efficient sampling strategy based on the distribution of RR sets, which can further reduce the computation time for typical social networks with long-tail distributions. Considering a NUMA architecture, we scale up our solution on 128-core CPUs and reduce the memory footprint by up to 45.7% with negligible time overhead (or even faster) and without perceivable loss of accuracy.

Chen, Xinyu↗

Molecular and Epidemiological Investigation of Fluconazole-resistant Candida parapsilosis —Georgia, United States, 2021

Abstract Background Reports of fluconazole-resistant Candida parapsilosis bloodstream infections are increasing. We describe a cluster of fluconazole-resistant C parapsilosis bloodstream infections identified in 2021 on routine surveillance by the Georgia Emerging Infections Program in conjunction with the Centers for Disease Control and Prevention. Methods Whole-genome sequencing was used to analyze C parapsilosis bloodstream infections isolates. Epidemiological data were obtained from medical records. A social network analysis was conducted using Georgia Hospital Discharge Data. Results Twenty fluconazole-resistant isolates were identified in 2021, representing the largest proportion (34%) of fluconazole-resistant C parapsilosis bloodstream infections identified in Georgia since surveillance began in 2008. All resistant isolates were closely genetically related and contained the Y132F mutation in the ERG11 gene. Patients with fluconazole-resistant isolates were more likely to have resided at long-term acute care hospitals compared with patients with susceptible isolates (P = .01). There was a trend toward increased mechanical ventilation and prior azole use in patients with fluconazole-resistant isolates. Social network analysis revealed that patients with fluconazole-resistant isolates interfaced with a distinct set of healthcare facilities centered around 2 long-term acute care hospitals compared with patients with susceptible isolates. Conclusions Whole-genome sequencing results showing that fluconazole-resistant C parapsilosis isolates from Georgia surveillance demonstrated low genetic diversity compared with susceptible isolates and their association with a facility network centered around 2 long-term acute care hospitals suggests clonal spread of fluconazole-resistant C parapsilosis. Further studies are needed to better understand the sudden emergence and transmission of fluconazole-resistant C parapsilosis.

Misas, Elizabeth (ORCID:0000000162437716)↗

Visual Analytics of Multivariate Networks With Representation Learning and Composite Variable Construction

Multivariate networks are commonly found in real-world data-driven applications. Uncovering and understanding the relations of interest in multivariate networks is not a trivial task. This article presents a visual analytics workflow for studying multivariate networks to extract associations between different structural and semantic characteristics of the networks (e.g., what are the combinations of attributes largely relating to the density of a social network?). The workflow consists of a neural-network-based learning phase to classify the data based on the chosen input and output attributes, a dimensionality reduction and optimization phase to produce a simplified set of results for examination, and finally an interpreting phase conducted by the user through an interactive visualization interface. A key part of our design is a composite variable construction step that remodels nonlinear features obtained by neural networks into linear features that are intuitive to interpret. We demonstrate the capabilities of this workflow with multiple case studies on networks derived from social media usage and also evaluate the workflow with qualitative feedback from experts.

97 MATHEMATICS AND COMPUTING↗

Evaluation of Graph Analytics Frameworks Using the GAP Benchmark Suite

The analysis of connected data is an increasingly important application in high-performance computing. Such analyses can reveal fraudulent patterns in financial transactions, optimize telecommunications networks, predict information flow in social networks, etc. However, the landscape of graph analytics is highly diverse. Graph algorithms stress processor architectures differently, and no one graph can represent all topologies. Consequently, no single approach or framework is expected to be optimal for all graph analytics problems. To help make sense of this diverse landscape, we evaluated four approaches to graph analytics: GraphBLAS, Galois, BGL17, GraphIt; and compare them against hand-tuned implementations that take advantage of hardware features on our test platform. Graph- BLAS formulates graph analytics as sparse linear algebra. Galois provides syntactic constructs for data parallelism over irregular data structures. BGL17 is a generic C++ template library for implementing graph algorithms. GraphIt provides a domain- specific language to describe and optimize graph algorithms. We use the GAP Benchmark Suite to establish baseline performance and guide the side-by-side evaluation of each framework. GAP consists of 30 tests: six graph analytics algorithms (breadth- first search, single-source shortest path, PageRank, betweenness centrality, connected components, and triangle counting) run on five graphs, each with different topological characteristics (e.g., high diameter, skewed degree distribution, high average degree). High-performance reference implementations are included for each benchmark algorithm. Because a graph can be loaded into memory a number of ways (e.g., flat file on disk, compressed sparse format, data frames, retrieved from SQL or NoSQL databases), our evaluation focused on computational performance rather than I/O. Our results show the relative strengths of each framework.

Graph algorithms, Benchmarking, shared-memory prog↗

Communications Dashboard (Control Rooms Take a Cue from Facebook), Chapter 1

Papers published via IEEE and AIAA conferences have presented an overview of how social media could benefit NASA working environments in general and proposed three specific social applications to benefit space flight control operations. One of them, Communications Dashboard, would help a real time flight controller keep up with both the "big picture" and significant details of operations via a cohesive interface similar to those of social networking services (SNS). Instead of recreational social features, "CommDash" would support functions like console logging, categorized and threaded text chat streams with enhanced accountability and graphics display features, high-level status displays driven by telemetry or other events, and an on-screen hailing function for requesting voice or text stream conversation. Moving certain voice conversations to text streams would reduce confusion and stress in two ways. Within text conversations, there would be far less repetition of content since text conversations have visual persistence and are reviewable instantly, e.g., there s no need to brief new participants to a discussion -- they just read what s already there. Remaining voice traffic would stand out more clearly, and quieter voice loops means fewer "say again" calls and less distraction from visual and mental tasks, thus less stress. (Most flight controllers monitor 4 or 5 voice loops at once.) Links could be created from console log entries to chat selections so that underlying details are readily available yet unobtrusive. This would reduce the confusion that rises from having multiple and sometimes divergent copies of the same information due to cut/copy and paste operations, attachments, and asynchronous editing. This concept could apply to a plethora of real time control environments and to other settings with lots of information juggling. This paper explores the dashboard concept in further detail and chronicles the first phase of a NASA IT Labs (Information Technology) project that could lead to a working system

Scott, David w.↗

SpaceOps 2012 Plus 2: Social Tools to Simplify ISS Flight Control Communications and Log Keeping

A paper written for the SpaceOps 2012 Conference (Simplify ISS Flight Control Communications and Log Keeping via Social Tools and Techniques) identified three innovative concepts for real time flight control communications tools based on social mechanisms: a) Console Log Tool (CoLT) - A log keeping application at Marshall Space Flight Center's (MSFC) Payload Operations Integration Center (POIC) that provides "anywhere" access, comment and notifications features similar to those found in Social Networking Systems (SNS), b) Cross-Log Communication via Social Techniques - A concept from Johnsson Space Center's (JSC) Mission Control Center Houston (MCC-H) that would use microblogging's @tag and #tag protocols to make information/requests visible and/or discoverable in logs owned by @Destination addressees, and c) Communications Dashboard (CommDash) - A MSFC concept for a Facebook-like interface to visually integrate and manage basic console log content, text chat streams analogous to voice loops, text chat streams dedicated to particular conversations, generic and position-specific status displays/streams, and a graphically based hailing display. CoLT was deployed operationally at nearly the same time as SpaceOps 2012, the Cross- Log Communications idea is currently waiting for a champion to carry it forward, and CommDash was approved as a NASA Iinformation Technoloby (IT) Labs project. This paper discusses lessons learned from two years of actual CoLT operations, updates CommDash prototype development status, and discusses potential for using Cross-Log Communications in both MCC-H and/or POIC environments, and considers other ways for synergizing console applcations.

Cowart, Hugh S.↗

Scalable edge clustering of dynamic graphs via weighted line graphs

Timestamped relational datasets consisting of records (or connections) between pairs of entities are ubiquitous in network science. For applications like peer-to-peer communication, email, various social network interactions, and computer network security, it is useful to organize these records into groups based on how and when they are occurring. Weighted line graphs offer a natural way to model how records are related in such datasets but for large real-world graph topologies, building and utilizing the line graph is prohibitively expensive. Here, we present the framework to cluster the edges of a dynamic graph via the associated line graph that contains two major contributions. The first is a method to work with the line graph implicitly and the second is a distributed scale implementation of an agglomerative hierarchical graph clustering algorithm. We outline a novel hierarchical dynamic graph edge clustering approach that efficiently breaks massive relational datasets into small sets of edges containing events at various timescales. This is in stark contrast to traditional graph clustering algorithms that prioritize highly connected (clique-like) community structures. Our approach relies on constructing a sufficient subgraph of a weighted line graph and applying a hierarchical agglomerative clustering. This approach is related to scalable techniques from spatial clustering, nonlinear-dimension reduction, topological data analysis, and draws particular inspiration from HDBSCAN. As an edge clustering, this method yields an overlapping node clustering. Our algorithm is parallelizable and we demonstrate efficient clustering of a billion-scale, real-world dynamic graph into small edge sets that correlate in topology and time. The entire clustering process for a graph with tens of billions of edges takes just a few minutes of run time on 256 nodes of a distributed compute environment. We argue how the output of the edge clustering is useful for a multitude of data visualization and powerful machine learning tasks, both involving the original massive dynamic graph data and metadata associated with the nodes and edges. Finally, we describe how this approach can be extended to dynamic hypergraphs and dynamic graphs/hypergraphs with unstructured data living on vertices and edges.

Data Analysis↗