Ensuring Scientific Reproducibility within the Earth Observation Community: Standardized Algorithm Documentation for Improved Scientific Data Understanding
Explore the source record for details and available documents.
SEARCH · Engineering Papers
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
NASA’s Life Sciences Ports (NLSP) serves the scientific community by providing curated data from space life science experiment. The Human Research Program (HRP) with the help of NLSP is currently transforming their life sciences data archive systems and processes to improve compliance with the FAIR principles [1]. Some of these improvements will at the same time support the twin pillars of Open Science [2]: transparency of methods and reproducibility of results. Scientific transparency is marked by the easily intelligible communication of what has been investigated: what were the procedures for collecting sample and the characteristics of samples collected? what kinds of measurements were made, what were the environmental conditions of the measurements? What were the analysis techniques of the collected data? Reproducibility of the results and findings from the investigation requires a high level of transparency for all but the simplest investigations; the slightest deviation in communicating and replicating complex experimental procedures or data analyses can often yield quite different data and even findings, thwarting their validation. One of the ways the NLSP is aiming to improve the communication of scientific information is through the use of ontology-driven metadata. Ontologies are powerful, graph-based knowledge representation structures, which can be leveraged to increase data interoperability, the area of the FAIR principles in which many data systems most lack compliance. Over the past decade, there has been a concerted effort in the biomedical community to develop modular and narrowly focused domain and application-specific ontologies in a common, open-source framework, the Open Biological and Biomedical Ontology (OBO) Foundry [3]. The open sharing and modular nature of this effort promises huge increases in harmonized data sharing for systems that leverage these models. Which is in line with the FAIR Data Principles of Findability, Accessibility, Interoperability, and Reuse for scientific data management and stewardship. 1. Wilkinson, M.D., et al., The FAIR Guiding Principles for scientific data management and stewardship. Sci Data, 2016. 3: p. 160018. 2. National Academies of Sciences, E. and Medicine, Open Science by Design: Realizing a Vision for 21st Century Research. 2018, Washington, DC: The National Academies Press. 232. 3. Smith, B., et al., The OBO Foundry: coordinated evolution of ontologies to support biomedical data integration. Nat Biotechnol, 2007. 25(11): p. 1251-5.
NASA’s Life Sciences Ports (NLSP) serves the scientific community by providing curated data from space life science experiment. The Human Research Program (HRP) with the help of NLSP is currently transforming their life sciences data archive systems and processes to improve compliance with the FAIR principles [1]. Some of these improvements will at the same time support the twin pillars of Open Science [2]: transparency of methods and reproducibility of results. Scientific transparency is marked by the easily intelligible communication of what has been investigated: what were the procedures for collecting sample and the characteristics of samples collected? what kinds of measurements were made, what were the environmental conditions of the measurements? What were the analysis techniques of the collected data? Reproducibility of the results and findings from the investigation requires a high level of transparency for all but the simplest investigations; the slightest deviation in communicating and replicating complex experimental procedures or data analyses can often yield quite different data and even findings, thwarting their validation. One of the ways the NLSP is aiming to improve the communication of scientific information is through the use of ontology-driven metadata. Ontologies are powerful, graph-based knowledge representation structures, which can be leveraged to increase data interoperability, the area of the FAIR principles in which many data systems most lack compliance. Over the past decade, there has been a concerted effort in the biomedical community to develop modular and narrowly focused domain and application-specific ontologies in a common, open-source framework, the Open Biological and Biomedical Ontology (OBO) Foundry [3]. The open sharing and modular nature of this effort promises huge increases in harmonized data sharing for systems that leverage these models. Which is in line with the FAIR Data Principles of Findability, Accessibility, Interoperability, and Reuse for scientific data management and stewardship.
NASA's traditional science data processing systems have focused on specific missions, and providing data access, processing and services to the funded science teams of those specific missions. Recently NASA has been modifying this stance, changing the focus from Missions to Measurements. Where a specific Mission has a discrete beginning and end, the Measurement considers long term data continuity across multiple missions. Total Column Ozone, a critical measurement of atmospheric composition, has been monitored for'decades on a series of Total Ozone Mapping Spectrometer (TOMS) instruments. Some important European missions also monitor ozone, including the Global Ozone Monitoring Experiment (GOME) and SCIAMACHY. With the U.S.IEuropean cooperative launch of the Dutch Ozone Monitoring Instrument (OMI) on NASA Aura satellite, and the GOME-2 instrumental on MetOp, the ozone monitoring record has been further extended. In conjunction with the U.S. Department of Defense (DoD) and the National Oceanic and Atmospheric Administration (NOAA), NASA is now preparing to evaluate data and algorithms for the next generation Ozone Mapping and Profiler Suite (OMPS) which will launch on the National Polar-orbiting Operational Environmental Satellite System (NPOESS) Preparatory Project (NPP) in 2010. NASA is constructing the Science Data Segment (SDS) which is comprised of several elements to evaluate the various NPP data products and algorithms. The NPP SDS Ozone Product Evaluation and Test Element (PEATE) will build on the heritage of the TOMS and OM1 mission based processing systems. The overall measurement based system that will encompass these efforts is the Atmospheric Composition Processing System (ACPS). We have extended the system to include access to publically available data sets from other instruments where feasible, including non-NASA missions as appropriate. The heritage system was largely monolithic providing a very controlled processing flow from data.ingest of satellite data to the ultimate archive of specific operational data products. The ACPS allows more open access with standard protocols including HTTP, SOAPIXML, RSS and various REST incarnations. External entities can be granted access to various modules within the system, including an extended data archive, metadata searching, production planning and processing. Data access is provided with very fine grained access control. It is possible to easily designate certain datasets as being available to the public, or restricted to groups of researchers, or limited strictly to the originator. This can be used, for example, to release one's best validated data to the public, but restrict the "new version" of data processed with a new, unproven algorithm until it is ready. Similarly, the system can provide access to algorithms, both as modifiable source code (where possible) and fully integrated executable Algorithm Plugin Packages (APPs). This enables researchers to download publically released versions of the processing algorithms and easily reproduce the processing remotely, while interacting with the ACPS. The algorithms can be modified allowing better experimentation and rapid improvement. The modified algorithms can be easily integrated back into the production system for large scale bulk processing to evaluate improvements. The system includes complete provenance tracking of algorithms, data and the entire processing environment. The origin of any data or algorithms is recorded and the entire history of the processing chains are stored such that a researcher can understand the entire data flow. Provenance is captured in a form suitable for the system to guarantee scientific reproducability of any data product it distributes even in cases where the physical data products themselves have been deleted due to space constraints. We are currently working on Semantic Web ontologies for representing the various provenance information. A new web site focusing on consolidating informaon about the measurement, processing system, and data access has been established to encourage interaction with the overall scientific community. We will describe the system, its data processing capabilities, and the methods the community can use to interact with the standard interfaces of the system.
This presentation provides the motivation for and status of implementation of persistent identifiers in NASA's Earth Observation System Data and Information System (EOSDIS). The motivation is provided from the point of view of long-term preservation of datasets such that a number of questions raised by current and future users can be answered easily and precisely. A number of artifacts need to be preserved along with datasets to make this possible, especially when the authors of datasets are no longer available to address users questions. The artifacts and datasets need to be uniquely and persistently identified and linked with each other for full traceability, understandability and scientific reproducibility. Current work in the Earth Science Data and Information System (ESDIS) Project and the Distributed Active Archive Centers (DAACs) in assigning Digital Object Identifiers (DOI) is discussed as well as challenges that remain to be addressed in the future.
Algorithm Theoretical Basis Documents (ATBDs) are documents which accompany Earth observation data products generated from algorithms. While ATBDs are essential to scientific reproducibility, these key documents are not standardized and are often difficult to find. In this paper, we present the prototype Algorithm Publication Tool (APT), a cloud-based ATBD authoring and editing tool for NASA’s Earth science data systems. A standardized ATBD information model is also described as well as lessons learned from developing the prototype tool.
The National Aeronautic and Space Administration (NASA) Commercial Smallsat Data Acquisition (CSDA) Program was established to identify, evaluate, and acquire data from commercial satellite companies that support and complement NASA’s Earth sciences missions and research goals. Since becoming a sustained program in 2020, CSDA has been developing a data system to provide scalable, efficient, continuous, and repeatable data management processes for all commercial data acquired by NASA. This includes providing access to commercial data to approved science investigators through NASA developed and vendor operated interfaces. As the commercial data user community and volume of data acquired by CSDA has grown, the team managing these data has adapted technologies and services to better address the data and information search, discovery, and access needs of the science user community. This presentation will provide an overview of CSDA Program data system activities including data availability, development of CSDA user interfaces, challenges in managing large-volume diverse datasets, and long-term data preservation activities to retain data for scientific reproducibility.
Reproducibility of scientific research relies on accurate and precise citation of data and the provenance of that data. Earth science data are often the result of applying complex data transformation and analysis workflows to vast quantities of data. Provenance information of data processing is used for a variety of purposes, including understanding the process and auditing as well as reproducibility. Certain provenance information is essential for producing scientifically equivalent data. Capturing and representing that provenance information and assigning identifiers suitable for precisely distinguishing data granules and datasets is needed for accurate comparisons. This paper discusses scientific equivalence and essential provenance for scientific reproducibility. We use the example of an operational earth science data processing system to illustrate the application of the technique of cascading digital signatures or hash chains to precisely identify sets of granules and as provenance equivalence identifiers to distinguish data made in an an equivalent manner.
The infrastructure of the Heliophysics discipline has promising components but with several missing gaps, drastically reducing research and development efficiency. Developing an online discovery and analysis software ecosystem will close several of these gaps. The five main focuses on this ecosystem should be Discovery, Implementation, Analysis, Reproducibility, and Sharing of results (DIARieS). In this paper, we give a detailed description of how the proposed software ecosystem should operate, and point out the large range of possible applications to benefit many disparate groups, such as researchers, operational staff, decision-makers, and educators. The infrastructure components and technological capabilities necessary for its completion are either currently available or in development, making such an ecosystem possible for the first time. One main focus of current infrastructure investments must be to adapt and connect these pieces together into a cohesive whole to increase our research and development efficiency.
The National Climate Assessment of the U.S. Global Change Research Program (USGCRP) analyzes and presents the impacts of climate change on the United States. The provenance information in the assessment is important because the assessment findings are of great public and academic concern and are used in policy and decision-making. By applying a use case-driven iterative methodology, we developed information models and ontology to represent the content structure of the recent National Climate Assessment draft report and its associated provenance information. We tested the ontology by using it in pilot systems serving information about instances of chapters, scientific findings, figures, tables, images, datasets, references, people, and organizations, etc. in the draft report, as well as interrelationships among those instances. The results successfully help users trace provenance in the draft report, such as finding all the journal articles from which a figure in the report was derived. The provenance information in our work was maintained in the context of the "Web of Data". In addition to the pilot systems we developed, other tools and services are also able to retrieve and utilize the provenance information. Our work is part of a Global Change Information System coordinated by the USGCRP that will eventually cover provenance information for the entire scope of global change research. Such a system will greatly increase understanding, credibility and trust in the global change research and foster reproducibility of scientific results and conclusions.
NASA Earth Exchange (NEX) is a data, supercomputing and knowledge collaboratory that houses NASA satellite, climate and ancillary data where a focused community can come together to address large-scale challenges in Earth sciences. As NEX has been growing into a petabyte-size platform for analysis, experiments and data production, it has been increasingly important to enable users to easily retrace their steps, identify what datasets were produced by which process chains, and give them ability to readily reproduce their results. This can be a tedious and difficult task even for a small project, but is almost impossible on large processing pipelines. We have developed an initial reproducibility and knowledge capture solution for the NEX, however, if users want to move the code to another system, whether it is their home institution cluster, laptop or the cloud, they have to find, build and install all the required dependencies that would run their code. This can be a very tedious and tricky process and is a big impediment to moving code to data and reproducibility outside the original system. The NEX team has tried to assist users who wanted to move their code into OpenNEX on Amazon cloud by creating custom virtual machines with all the software and dependencies installed, but this, while solving some of the issues, creates a new bottleneck that requires the NEX team to be involved with any new request, updates to virtual machines and general maintenance support. In this presentation, we will describe a solution that integrates NEX and Docker to bridge the gap in code-to-data migration. The core of the solution is saemi-automatic conversion of science codes, tools and services that are already tracked and described in the NEX provenance system, to Docker - an open-source Linux container software. Docker is available on most computer platforms, easy to install and capable of seamlessly creating and/or executing any application packaged in the appropriate format. We believe this is an important step towards seamless process deployment in heterogeneous environments that will enhance community access to NASA data and tools in a scalable way, promote software reuse, and improve reproducibility of scientific results.
The FAIR principle (findable, accessible, interoperable, and reusable) governs the storage and sharing of NASA space biology and health data[1]. These guiding principles maximize reuse of data and the reproducibility of scientific findings. The NASA Open Science Data Repository (OSDR; an expansion of NASA GeneLab) was built on the FAIR principles and houses over 500 studies and close to 1000 datasets from decades of space life sciences experiments. OSDR embodies the FAIR principles through data governance that includes mediated, embargoed, and fully open access data. The FAIR data governance principles were recently proposed to be expanded to encompass a FAIREST framework for assessing research data repositories (FAIR + Engagement, Social connections, and Trust)[2]. FAIREST emphasizes the importance of data repositories engaging with the scientific community and gaining the trust of researchers regarding data quality. Trust also refers to the TRUST principles developed for assessment of digital repositories: Transparency, Responsibility, User Focus, Sustainability, Technology[3]. We present the “Open Science for Life in Space” Analysis Working Groups (AWGs) as evidence regarding the power of engagement, social connections, and trust which has enhanced OSDR’s capabilities and productivity. AWG members engage in two main activities. One, members provide feedback on OSDR scientific standards for data ingestion, curation, and reuse (study, subject and assay metadata; processing pipelines; dataset formats and uniformed structures for machine-readability). Two, AWG members collaborate to mine-reuse OSDR data to conduct scientific analysis. With nearly 800 active members, the AWGs have resulted in 32 publications re-using OSDR data and contributed many papers in two major special issues in Cell (2020) and Nature (2024). AWGs also serve as networking groups, facilitate social connections between researchers at all levels of experience, and also have a social online ‘Forum’ used to keep members informed on projects and opportunities. This community-centric, productive, and trustworthy data culture has resulted in a broader effect with international space agencies, academics, and the commercial space sector wanting to submit their data to OSDR. Ten studies of Inspiration 4 data were recently publicly released by OSDR, as were some JAXA human data. Coming up soon in OSDR are data submissions from the European Space Agency, Virgin Galactic PIs, and SpaceX Polaris Dawn. A major benefit of OSDR is the array of standardized and uniformly formatted data (which was developed through AWG member consensus), from which visualization tools, analysis tools, and machine learning models can be built or trained. This talk will cover the Multi-Study Visualization Tool, the Environmental Data Application, RadLab, and a UCSF-NSF funded knowledge graph biomedical health discovery tool ‘SPOKE’ currently being integrated with OSDR. OSDR also provides training programs in bioinformatics and machine learning to improve the scientific community’s awareness of data availability and to boost their ability to perform data analysis. The increasing engagement of the scientific community and the public with technologies powered by artificial intelligence (AI) heightens the need for data analysis to be transparent. The AI for Life in Space initiative leverages the data products provided in OSDR to train AI models, with an emphasis on explainable and trustworthy AI, which would not be possible without FAIR data and metadata. Overall, here we will demonstrate the importance for NASA life sciences data repositories to adhere to the FAIREST framework, by providing examples and success stories from different aspects of OSDR.
Information quality is of paramount importance to science. Accurate, scientifically vetted and statistically meaningful and, ideally, reproducible information engenders scientific trust and research opportunities. Therefore, so-called Highly Influential Scientific Assessments (HISA) such as the U.S. Third National Climate Assessment undergo a very rigorous process to ensure transparency and credibility. As an activity to support the transparency of such reports, the U.S. Global Change Research Program has developed the Global Change Information System (GCIS). Specifically related to the transparency of NCA3, a recent activity was carried out to trace the provenance as completely as possible for all figures in the NCA3 report that predominantly used NASA data. This paper discusses lessons learned from this activity that trace the provenance of NASA figures in a major HISA-class pdf report.
Information quality is of paramount importance to science. Accurate, scientifically vetted and statistically meaningful and, ideally, reproducible information engenders scientific trust and research opportunities. Not surprisingly, federal bodies (e.g., NASA, NOAA, USGS) have very strictly affirmed the importance of information quality in their product requirements. So-called Highly Influential Scientific Assessments (HISA) such as The Third US National Climate Assessment (NCA3) published in 2014 undergo a very rigorous review process to ensure transparency and credibility. To support the transparency of such reports, the U.S. Global Change Research Program (USGCRP) has developed the Global Change Information System (GCIS). A recent activity was performed to trace the provenance as completely as possible for all NCA3 figures that were predominantly based on NASA data. This poster presents the mechanics of that project and the lessons learned from that activity.
Explore the source record for details and available documents.
The 2015 Space Radiation Standing Review Panel (from here on referred to as the SRP) met for a site visit in Houston, TX on December 8 - 9, 2015. The SRP met with representatives from the Space Radiation Element and members of the Human Research Program (HRP) to review the updated research plan for the Risk of Radiation Carcinogenesis Cancer Risk. The SRP also reviewed the newly revised Evidence Reports for the Risk of Acute Radiation Syndromes Due to Solar Particle Events (SPEs) (Acute Risk), the Risk of Acute (In-flight) and Late Central Nervous System Effects from Radiation Exposure (CNS Risk), and the Risk of Cardiovascular Disease and Other Degenerative Tissue Effects from Radiation (Degen Risk), as well as a status update on these Risks. The SRP would like to commend Dr. Simonsen, Dr. Huff, Dr. Nelson, and Dr. Patel for their detailed presentations. The Space Radiation Element did a great job presenting a very large volume of material. The SRP considers it to be a strong program that is well-organized, well-coordinated and generates valuable data. The SRP commended the tissue sharing protocols, working groups, systems biology analysis, and standardization of models. In several of the discussed areas the SRP suggested improvements of the research plans in the future. These include the following: It is important that the team has expanded efforts examining immunology and inflammation as important components of the space radiation biological response. This is an overarching and important focus that is likely to apply to all aspects of the program including acute, CVD, CNS, cancer and others. Given that the area of immunology/inflammation is highly complex (and especially so as it relates to radiation), it warrants the expansion of investigators expertise in immunology and inflammation to work with the individual research projects and also the NASA Specialized Center of Research (NSCORs). Historical data on radiation injury to be entered into the Watson “big data” study must be used with caution. The general scientific issues of reproducibility, details of experimental methods and data analysis from preclinical and basic research laboratories have been raised broadly over the last few years (not specific to this work) and indicate that caution must be applied in the ways these data are used. This pertains to preclinical data and also to phase 3 clinical trials in radiation oncology and medical oncology. Of course, appropriate use and analysis of these “big-data” sets also offer the potential of pinpointing limitations and extracting remaining useful information. Emphasis should be placed on the latter possibility. A key target is risk reduction from radiation exposure. Progress of the entire space program, now moving towards the Mars mission, requires timely answers to key components of human risk, which are known to be complex. Periodic review of progress should be conducted with additional resources directed into achieving critical milestones. Turning the long red bars to yellow and green (or for some risks such as CNS possibly to grey) must be high priority. That such progress will require new science and not engineering means that it should be viewed in a knowledge-based light. The technology-based aspects of engineering issues are certainly as important, however, science and knowledge-based problems are solved in a different way than engineering. Timelines for engineering are more predictable, while for science, progress can be methodical with occasional major incremental findings that can rapidly change the rate of progress. As opportunities for rapid incremental changes arise, periodic enhancement of investment is strongly recommended to enable such new knowledge to be quickly and efficiently exploited. Collaborations and linkages with National Institute of Allergy and Infectious Diseases (NIAID), the Biomedical Advanced Research and Development Authority (BARDA) and the Department of Defense (DoD) are in place and more are encouraged, where possible, with the radiation injury and medical countermeasure studies. This could include utilizing some of their animal model testing contracts to facilitate obtaining results using common platforms. Such approach will facilitate the comparison of results among laboratories, and will facilitate and accelerate the development of medical countermeasures. It is particularly noteworthy that the NASA Space Radiation Element is reaching out to the Multidisciplinary European Low Dose Initiative (MELODI) platform coordinating low dose radiation risk research, and to other international agencies that are studying low dose radiation effects in an effort to fill the void generated by the cancelation of the Department of Energy (DOE) low dose radiation program. While NASA is working actively with NIAID and BARDA to integrate their relevant findings of radiation mitigator investigations to NASA programs, the committee notes its disappointment that the United States currently lacks a dedicated low dose radiation program with clear mechanistic orientation and aimed at the quantification and mitigation of human radiation risk on Earth. This void gives to the NASA Space Radiation Program Element special societal value, but also makes its overall design more challenging.
There is currently too little reproducible data for a scientifically valid understanding of the initial responses of a diverse human population to weightlessness and other space flight factors. Astronauts on orbital space flights to date have been extremely healthy and fit, unlike the general human population. Data collection opportunities during the earliest phases of space flights to date, when the most dynamic responses may occur in response to abrupt transitions in acceleration loads, have been limited by operational restrictions on our ability to encumber the astronauts with even minimal monitoring instrumentation. The era of commercial personal suborbital space flights promises the availability of a large (perhaps hundreds per year), diverse population of potential participants with a vested interest in their own responses to space flight factors, and a number of flight providers interested in documenting and demonstrating the attractiveness and safety of the experience they are offering. Voluntary participation by even a fraction of the flying population in a uniform set of unobtrusive biomedical data collections would provide a database enabling statistical analyses of a variety of acute responses to a standardized space flight environment. This will benefit both the space life sciences discipline and the general state of human knowledge.
Among the primary objectives of the Open/Open-Source Science paradigm are making scientific investigation data transparent and results reproducible [1], objectives shared by the FAIR principles [2]. To accomplish this, the conceptual framework that includes all the investigation objects needs to be accurately captured and communicated to all data consumers. A large part of this requires using metadata standards to annotate data collected. These standards should be readily accessible, informed by scientific community consensus and sufficiently specific to encompass all of the important aspects of the investigation. Starting in 2020 we have been co-leading an open consortium to develop a new metadata standard, the Radiation Biology Ontology (RBO), through the Open Biological and Biomedical Ontologies (OBO) Foundry [3]. We began by transforming many of the terms from the National Council on Radiation Protection and Measurement into concepts that can be formally related to existing OBO Foundry classes or attributes. We then identified and imported into the RBO existing OBO Foundry classes that have obvious relevance for radiation biomedicine (for example, concepts from the Environment Ontology that describe radiative processes, and concepts from the Gene Ontology dealing with molecular and cellular responses to radiation). Finally, we scrutinized datasets from investigations of radiation effects held in NASA GeneLab and LSDA repositories and added additional classes, instances, and attributes into the RBO that should be used to annotate these data. We developed the RBO using the open-source tools of GitHub and publish the RBO periodically through the NIH/NCBI BioPortal website, so systems worldwide can leverage the knowledge it contains [4]. This initial phase of concept modeling has yielded an RBO that at present has more than 300 declared concepts, with more than 3500 additional concepts imported from other OBO Foundry ontologies. While this first phase has focused on concepts for annotating samples, environments, exposures, and measurements, the next phase will center on supporting annotation of results and findings, such as concept models of molecular, cellular and tissue effects. The value of the RBO will be determined in part by our ability to engage the community in its development, and we have established a Radiobiology Informatics Consortium with unrestricted membership as the owner of the RBO in order to encourage investigators, system owners and other to join in this effort. Anyone can report issues or request new concept modeling or other features directly on GitHub. By using the BioPortal application programming interface, systems can pose dynamic queries to the latest version of the RBO for information on individual classes or entire hierarchies; this design eliminates the need for systems to be updated in order to use newer versions of the RBO. We hope to contribute to the advancement of open radiobiological science through the continued, open development of the RBO, that will provide more precise, machine-interpretable descriptions of investigations, as well as support data meta-analysis through machine learning or other artificial intelligence methods. REFERENCES [1] Open science in space. Nature Medicine, 2021. 27(9): p. 1485-1485. [2] Wilkinson, M.D., et al., The FAIR Guiding Principles for scientific data management and stewardship. Sci Data, 2016. 3: p. 160018. [3] Smith, B., et al., The OBO Foundry: coordinated evolution of ontologies to support biomedical data integration. Nat Biotechnol, 2007. 25(11): p. 1251-5. [4] Whetzel, P.L., et al., BioPortal: enhanced functionality via new Web services from the National Center for Biomedical Ontology to access and use ontologies in software applications. Nucleic Acids Res, 2011. 39(Web Server issue): p. W541-5.