The ICGA Data Visualization Platform

The Indian Cancer Genome Atlas (ICGA) Portal is the official data visualization and exploration platform of the Indian Cancer Genome Atlas.

Built on the open-source cBioPortal framework, the Portal hosts curated, harmonized, clinically annotated, multi-omics cancer datasets generated through ICGA studies. It enables researchers to explore genomic, transcriptomic, proteomic, and associated clinical datasets through an intuitive, browser-based interface while operating within ICGA’s governed data access framework.

The Portal currently hosts datasets from the ICGA Breast Cancer Programme and will expand to additional cancer types as new studies are completed.


About the ICGA Portal

The ICGA Portal provides researchers with secure, browser-based access to curated and processed cancer genomics datasets.

The Portal enables exploratory and hypothesis-driven research through interactive visualisation and analytical tools while ensuring compliance with applicable ethical approvals, institutional governance policies, and regulatory requirements.

The Portal visualises highly processed, curated, harmonised, secondary-level cancer genomics datasets. It is intended for research use and is not designed to host raw sequencing data.

The Portal has been developed with philanthropic support from Strand Life Sciences Ltd.

introduction to ICGA Data Portal

Data Available Through the Portal

The portal provides secondary-level processed datasets, including:

Genomic Data

  • Somatic variant data
  • Mutation Annotation Format (MAF)
  • Copy Number Alterations (CNA)
  • Mutation summaries

Transcriptomic Data

  • Processed gene expression matrices
  • Expression summaries

Proteomic Data

  • Processed protein summaries

Clinical Data

  • Demographic variables (where available)
  • Tumour characteristics
  • Clinical staging
  • Treatment information (where available)
  • Longitudinal follow-up (where available)

All datasets are de-identified and curated before release.

Data Availability and Ethical Compliance

ICGA adheres to strict ethical and regulatory standards in data sharing.

  • All breast cancer datasets available through the portal are de-identified and limited to somatic mutations only.
  • Germline data is neither analyzed nor shared. This restriction is a result of ethical approvals that ensure patient privacy and compliance with responsible data-use guidelines.
  • Data is generated and processed using standardized pipelines designed to ensure quality and compliance.
  • Access is governed by the ICGA Data Access Committee (DAC).

Current Study

ICGA Breast Cancer Cohort

The current release contains clinically annotated breast cancer datasets generated across participating ICGA centres.

The cohort includes:

  • Treatment-naïve breast cancer patients
  • Matched tumour and adjacent non-tumour samples
  • Patients aged 18–90 years
  • Multi-omics profiling including:
    • Whole Genome Sequencing (WGS)
    • Whole Exome Sequencing (WES)
    • RNA Sequencing
    • Proteomics
  • Harmonised clinical metadata

Additional cancer types will be incorporated in future releases.

Metadata of the ICGA’s cohort on breast cancer patients

Data Access

The ICGA follows a controlled access model governed by the Data Access Committee (DAC) in alignment with ICGA Data Policy and DBT PRIDE guidelines.

Digital Personal Data Protection (DPDP)

ICGA is committed to implementing a privacy-by-design governance framework consistent with the principles of the Digital Personal Data Protection Act, 2023 and the Digital Personal Data Protection Rules, 2025.

Accordingly:

  • only the minimum data necessary for the approved research purpose will be provided;
  • access is limited to the approved investigators;
  • all access is logged and auditable;
  • approved data must be stored and processed only on India-based infrastructure unless otherwise authorised by ICGA; and
  • all approved users must comply with the ICGA Data Policy and Data User Agreement.

Access Routes

ICGA data is accessible through two routes. Both routes require a completed application and approval by the DAC before access is granted.

Route A — ICGA Portal (cBioPortal)

Interactive, browser-based access to processed and visualised datasets within the secure ICGA portal environment. Please remember no raw data would be available here. Researchers can explore somatic mutation profiles, gene expression patterns, proteomics summaries, and associated clinical metadata using the portal’s built-in analysis tools. Data remains within ICGA’s secure infrastructure at all times. [The portal visualizes highly processed, curated, and harmonized multidimensional cancer genomics data. ]

Route B — AWS Controlled Access / Secure Computational Workspace

Programmatic access to data files — VCF files, MAF files, expression matrices, and processed proteomics outputs — for researchers requiring computational analysis beyond what the portal interface supports. Access is provided within ICGA-managed, India-based infrastructure. Applicants are responsible for all associated AWS infrastructure costs. Contact ICGA for details.

Both routes are subject to a single consolidated application reviewed by the DAC.

Note: Commercial and industry applications are subject to additional review including execution of a Commercial Data Licensing Agreement. Contact suveera@icga.co.in with details of your intent before submitting any data request.

Application Process

Researchers seeking access must submit a single consolidated application that includes:

  • Research proposal and objectives
  • Details of investigators and institutional affiliation
  • Data requirements (type, scope, and access route requested)
  • Data management and security plan
  • Timelines and expected outcomes
  • Institutional ethics approvals
  • Conflict of interest disclosures
  • Details of any international or inter-institutional collaborations

All applications are reviewed by the ICGA Data Access Committee (DAC). Incomplete applications will not be considered. Full and final approval is followed by the signing of a Data User Agreement (DUA) with ICGA.

Processed Data Access (Conditional)

In cases where there is clear scientific justification and upon DAC approval, researchers may be granted access to processed data files within an ICGA-approved secure compute environment. Data does not leave ICGA’s governed infrastructure unless the DAC exceptionally approves it. The modality of access will be determined by ICGA on a case-by-case basis following DAC’s review.

Such requests must demonstrate:

  • A scientific need that cannot be met through portal-based analysis
  • Adequate institutional data security measures, including encryption at rest and in transit
  • Confirmation that ICGA dataset and its derivatives will be stored and processed in India-based infrastructure
  • Compliance with ICGA’s Data Policy and Data User Agreement

Eligible file types include processed somatic variant data (VCF/MAF), normalised expression matrices, and proteomics outputs.

Data Not Available

The following are not available for external access at this stage:

  • Raw sequencing data (BAM, FASTQ)
  • Raw proteomics data
  • Germline variant data
  • Any data that may increase re-identification risk

Requests for the above will not be considered under the current policy framework.

Note:

Given the strict regulatory framework of the Digital Personal Data Protection (DPDP) Act, your application must comprehensively address the following:

  • Full and Transparent Disclosure: You must declare the exact purpose of your research and explicitly list all collaborations and associated studies where this data will be used. Providing incomplete disclosures or utilizing the data for undeclared purposes strictly violates data access protocols and DPDP compliance.
  • Data Minimization & Scientific Rationale: Under DPDP principles, we only provision the minimum data necessary for the stated research, not the entire dataset. You must request only what you specifically need. Please clearly outline your required patient volume, specific tumor subtypes, and the precise clinical metadata you are seeking. You must provide a strong scientific rationale for requesting specific clinical variables and treatment histories, particularly as this is classified as sensitive personal data under the law.
  • Cross-Border Compliance: If your research involves any cross-border data transfer or falls under an international collaboration, you must clearly outline this and address the resulting DPDP Act implications in your application.

Conditions of Use

All approved users must comply with the following conditions throughout the approved access period:

  1. Data must be used strictly for the approved research purpose. Any change in research purpose, methodology, or key personnel must be notified to ICGA within 30 days and may require a new application.
  2. No attempt to re-identify any individual from ICGA de-identified datasets is permitted. This prohibition applies to all users and all methods of analysis.
  3. Data must not be shared with parties not named in the approved application and Data User Agreement.
  4. All ICGA data must be stored, accessed, and processed exclusively on India-based infrastructure. Transfer to servers or compute environments outside India is not permitted without explicit written approval from ICGA DAC.
  5. Significant derived datasets generated from ICGA data — such as integrated multi-omics outputs or novel variant call sets — should be shared back with ICGA to enable cumulative scientific value for the community.
  6. ICGA must be acknowledged in all oral presentations, written disclosures, and publications resulting from analyses of ICGA data.

Suggested citation:“The results [published or shown] here are based, in whole or in part, on data generated by the Indian Cancer Genome Atlas (ICGA) Network: https://icga.in, https://icga.net.in

ICGA Data, Resources, and Materials

ICGA is dedicated to advancing cancer research through a rigorous, end-to-end process that involves:

  • Collecting diverse biospecimens and clinical metadata from partnering hospitals and clinical centres across India
  • Generating molecular analytes for detailed multi-omics characterisation
  • Applying standardised sequencing, proteomic, and imaging methods
  • Curating and annotating data to enable responsible, reproducible research
  • Providing accessible data to the research community through a governed access framework

Requests for Biological Samples and Materials

Due to legal and ethical considerations, ICGA is unable to accommodate requests for biological samples, analytes, or tissue materials. All cases within the ICGA programme have been consented exclusively for ICGA use, and the redistribution of materials to outside parties is prohibited. Additionally, the majority of tissue samples have been depleted through the multiple assays performed for ICGA research.


Frequently Asked Questions (FAQ)

Loader image

The portal provides access to processed, secondary-level data, including:

  • Clinical data
  • Somatic mutation data
  • Copy Number Alterations
  • Gene expression summaries
  • Proteomics summaries

Yes, provided your specific access request was approved with download. Data can be downloaded from:

  • Study pages
  • Query results
  • Visualization panels

Users can also define custom cohorts (“virtual studies”) using clinical or genomic filters and download the corresponding datasets.

No.

The Portal contains processed, secondary-level datasets only.

FASTQ, BAM, CRAM and other raw sequencing files are not available.


Access to controlled datasets requires application through the ICGA Data Access Committee (DAC).

This includes:

  • Raw RNA-seq data (including count-level data)
  • Variant Call Format (VCF) files
  • Other detailed datasets not available through the portal

ICGA Data Access Form

For queries, contact:

  • suveera@icga.co.in
  • data-access@icga.co.in

No. Since the portal provides processed, gene-level summarized data rather than raw count matrices, workflows requiring raw count reprocessing are not supported directly from portal downloads.

The portal supports:

  • Gene-level CNA downloads through the query interface
  • Segment-level copy number downloads through study-specific links

Users can query specific genes or define cohorts before exporting results.

No. The ICGA Portal is an independent instance of the cBioPortal platform and hosts ICGA-specific datasets only.

No. The interface is designed to be intuitive. However, familiarity with cBioPortal workflows may be helpful for advanced queries and cohort analyses.

Yes. Users can apply clinical and genomic filters to create custom cohorts (“virtual studies”) for downstream exploration and data export.

Users are required to acknowledge the Indian Cancer Genome Atlas (ICGA) in any publications, presentations, or outputs derived from ICGA data.

Suggested citation:

“The results published or shown here are based, in whole or in part, on data generated by the Indian Cancer Genome Atlas (ICGA) Network: https://icga.in and https://icga.net.in.”

Where applicable, users should also cite associated ICGA publications relevant to the dataset used.

For citation-related queries, contact:

  • suveera@icga.co.in
  • data-access@icga.co.in

Yes. ICGA datasets are periodically updated as new data is generated, processed, and curated.

Breast cancer study data, for example, is routinely updated.

Users should record:

  • Study identifier
  • Date of data access

for reproducibility and future reference.

For dataset version queries, contact:

  • suveera@icga.co.in
  • data-access@icga.co.in