What is HELPER?
Data sources for the HELPER.
Database construction pipeline.
Four modules in HELPER.
Contact us.
HELPER (Helicobacter Pylori Encyclopedia for Research) is a comprehensive online database for exploring Helicobacter pylori (H. pylori) pathogen from different aspects, including population genetic, genomic variation, virulence factors and antibiotic resistance. It is a user-friendly interface with facilities to produce maps and explore visualizations.
In HELPER, users can:
√ Browse and download the host information (including geographic location, collection data and host disease).
√ Download the H. pylori genomes by population or country easily.
√ Browse the SNP annotation and frequency in different population and different disease.
√ Browse and download the virulence factors types for each H. pylori specific strain.
√ Browse the distribution of virulence factors by population or country.
√ Browse the prevalence of genotypic resistant strains in different countries, populations and time periods.
√ Browse frequencies of the resistance mutations in both genotypic resistant strains and sensitive strains.
HELPER comprised 4067 H. pylori genomes in 76 different countries including 180 stains which were newly sequenced and 3887 genomes downloaded from the National Center for Biotechnology Information (NCBI) database.
Four modules are displayed on the homepage of our website including: “Population Structure”, “SNP”, “Virulence Factor” and “Antibiotic Resistance”. Users can easily browse in HELPER by clicking on the appropriate box query.
We used three different and complementary approaches, fineSTRUCTURE, ADMIXTURE and DAPC to define the population structure of our dataset. Click the “Population” botton at the top of the homepage and select the population of interest. Users can see the distribution of the population in the world map and get summary data for each isolate in a table format. Click the “Country” button just next to the “Population” box and select the country of interest. Users can visualize the population composition of the specific country in a pie chart form and get detail information for each isolate in this country in a table format.
We mapped and called variants for 4067 H. pylori isolates using snippy tool with the genome of strain 26695 as reference. There are totally 364,269 SNPs and 1534 genes in our database. “SNP” section involves three main segments, including “Variant”, “Gene” and “Region”. In “Variant”, users can search for the SNP variation of interest by using the anchor position on the 26695 genome. Then, users can access three detailed modules. First, “Annotation” which provides gene symbols where the SNP is located and molecular consequence based on protein annotation. Second, “Allele Frequency” presents the frequencies for reference and alternate alleles in different populations. Third, “Allele Frequency (Case / Control)” displays the frequencies for reference and alternate alleles in stains isolated from gastric cancer and non-gastric cancer separately. In “Gene”, users can search for the gene symbol of interest and then they can obtain a list of SNPs on this gene. In “Region”, users can obtain SNPs within the region by searching for the start and end positions.
We investigated the virulence of each isolate by genotyping the two most important virulence factors of H. pylori, CagA and VacA. For CagA, we determined the biological activity of CagA by the types and sequence of EPIYA (glutamic acid-proline-isoleucine-tyrosine-alanine) segment at the C-terminal region. For VacA, we determined the genotype of VacA by visually inspecting the corresponding primers. Users can obtain the structure and function of CagA and VacA in each country in the “Country” section, and even in each population or subpopulation in the “Population” and “Subpopulation” module.
This section in our database provides information on genotypic antibiotic resistance, determined based on mutations identified in a subset of isolates with phenotypic resistance data.
Among 687 isolates with phenotypic antibiotic resistance data:
115 newly sampled isolates were tested using the micro broth dilution method.
The remaining 572 isolates were sourced from the NCBI database, with susceptibility determined by Etest or agar dilution methods.
Resistance mutations were identified by comparing the mutation frequencies in phenotypically resistant and susceptible strains using a Chi-squared test with Bonferroni correction (P<0.05).
A strain was classified as genotypically resistant if it carried one or more of these validated resistance-associated mutations.
Users can explore:
The prevalence of genotypic resistant strains across different populations, countries, and time periods.
The frequencies of resistance mutations in resistant and sensitive strains.
Guangfu Jin, M.D., Ph.D.
Professor of Epidemiology
Department of Epidemiology, School of Public Health, Nanjing Medical University
101 Long-mian Avenue, Nanjing 211166, China
TEL (FAX): +86-25-8686-8397
Email: guangfujin@njmu.edu.cn