GSPlot is an interactive web application for analyzing and visualizing gene set enrichment results as a 2D map of related gene sets.
Live deployment: https://gsplot.cs.umb.edu
Gene Set Enrichment Analysis (GSEA) can produce long result tables with many overlapping gene sets, which makes interpretation difficult. GSPlot helps address this problem by converting enrichment results into an interactive map where related gene sets are placed near each other.
GSPlot allows users to provide ranked, thresholded, or scored gene input, run enrichment analysis, measure gene-set similarity, reduce the results into a 2D embedding, and explore the output through an interactive graph. The goal of the project is to make large enrichment results easier to compare, filter, cluster, and interpret.
Main features include:
- Gene set enrichment analysis from user-provided gene input.
- Support for ranked, thresholded, and scored input modes.
- Built-in support for human and mouse MSigDB-derived gene set resources.
- Support for custom uploaded gene set collections.
- Pairwise gene-set distance calculation using similarity measures such as Jaccard distance and overlap coefficient.
- Dimensionality reduction using UMAP, t-SNE, or Isomap.
- Interactive visualization with point selection, filtering, clustering, and result export.
- Optional cluster label generation support.
Important project files and folders include:
manage.py: Django project entry point.gsplot/settings.py: Django settings and environment configuration.graph/views.py: API endpoints and analysis workflow orchestration.graph/dataReduction.py: enrichment, distance calculation, and dimensionality-reduction logic.graph/static/resources/: local gene set resource files used by the application.requirements.txt: pinned Python dependencies..gitignore: ignored local files, environment files, and sensitive files.
GSPlot relies on scientific libraries like NumPy and SciPy[cite: 1]. To ensure these packages compile correctly from source, install the required system-level C and Fortran development headers before creating your virtual environment:
sudo apt-get update
sudo apt-get install build-essential pkg-config python3-dev gfortran libopenblas-dev liblapack-dev
Follow these steps to download the code and set up the Python environment.
1. Clone the repository
git clone https://github.com/PathwayAndDataAnalysis/gsplot.git
cd gsplot
2. Create and activate a virtual environment
python3 -m venv .venv
source .venv/bin/activate
(On Windows, use: .venv\Scripts\activate)
3. Install dependencies Upgrade your core build tools first to prevent metadata generation errors, then install the pinned project requirements:
pip install --upgrade pip setuptools wheel
pip install -r requirements.txt
4. Configure environment variables
GSPlot requires a Django secret key. Create a local .env file in the project root:
DJANGO_SECRET_KEY=replace-with-a-strong-secret
If optional cluster label generation is used, also configure the Gemini API key:
GEMINI_API_KEY=replace-with-your-api-key
Note: Do not commit .env files, API keys, tokens, or secret keys to the repository.
5. Run database migrations
python manage.py migrate
For local development and testing, use Django's built-in development server.
1. Start the server
python manage.py runserver
2. Access the application Open the local development site at:
http://127.0.0.1:8000
For a live deployment, do not use the development server. Use a production-grade WSGI server (Gunicorn) and a reverse proxy (Nginx) to handle traffic and serve static files securely.
1. Collect static files Gather all static assets into a single directory for Nginx to serve:
python manage.py collectstatic
2. Install Gunicorn
pip install gunicorn
3. Create a Gunicorn systemd service Ensure Gunicorn runs in the background by creating a service file:
sudo nano /etc/systemd/system/gsplot.service
Paste the following configuration (replace /path/to/gsplot and YOUR_LINUX_USERNAME with your actual server details):
[Unit]
Description=Gunicorn daemon for GSPlot
After=network.target
[Service]
User=YOUR_LINUX_USERNAME
Group=www-data
WorkingDirectory=/path/to/gsplot
ExecStart=/path/to/gsplot/.venv/bin/gunicorn --access-logfile - --workers 3 --bind 127.0.0.1:8000 gsplot.wsgi:application
[Install]
WantedBy=multi-user.target
Start and enable the service to run on boot:
sudo systemctl start gsplot
sudo systemctl enable gsplot
4. Configure Nginx as a reverse proxy Install Nginx and create a server block to route public traffic (port 80) to Gunicorn (port 8000):
sudo apt install nginx
sudo nano /etc/nginx/sites-available/gsplot
Paste the following configuration (update YOUR_SERVER_IP_OR_DOMAIN and /path/to/gsplot):
server {
listen 80;
server_name YOUR_SERVER_IP_OR_DOMAIN;
location = /favicon.ico { access_log off; log_not_found off; }
location /static/ {
alias /path/to/gsplot/staticfiles/;
}
location / {
include proxy_params;
proxy_pass http://127.0.0.1:8000;
}
}
Enable the site, check for syntax errors, and restart Nginx:
sudo ln -s /etc/nginx/sites-available/gsplot /etc/nginx/sites-enabled/
sudo nginx -t
sudo systemctl restart nginx
Grant Nginx read permissions to your home directory if necessary so it can access the staticfiles folder:
sudo chmod 755 /path/to/your/home/directory
The required Python packages are listed in requirements.txt.
Main dependencies include:
- Django
- numpy
- pandas
- scipy
- scikit-learn
- statsmodels
- umap-learn
- hdbscan
- gseapy
- python-dotenv
- google-genai and google-api-core for optional cluster label generation support
For reproducible results, use the pinned package versions in requirements.txt.
- Open the app.
- Choose an input mode:
- Ranked Genes
- Thresholded Genes
- Scored Genes
- Select the directional hypothesis where supported:
- Positive
- Negative
- Two-Sided
- Choose a gene set resource:
- Human MSigDB-derived resource
- Mouse MSigDB-derived resource
- Custom uploaded gene set resource
- Select one or more gene set collections.
- Set the minimum number of matched or relevant genes required for each gene set.
- Submit the analysis.
- Explore the interactive graph.
- Select points, inspect gene sets, apply clustering options, and export the results.
Ranked gene input accepts pasted text or uploaded .txt / .tsv files.
Each line should contain one gene symbol. If multiple columns are provided, the first column is used as the gene symbol.
Example:
TP53
MYC
STAT1
CXCL10
BRCA1
Thresholded input uses two gene lists:
- Significant genes
- Background or insignificant genes
Genes may be separated by commas, spaces, tabs, or new lines.
Example significant gene list:
TP53
MYC
STAT1
CXCL10
Example insignificant/background gene list:
GAPDH
ACTB
RPLP0
HPRT1
Scored input accepts uploaded .txt or .tsv files with exactly two columns and no header:
gene<TAB>score
Example:
TP53 2.84
MYC 1.91
STAT1 -1.32
CXCL10 -2.20
BRCA1 0.75
Users may upload custom gene set collections in supported formats such as:
.json.txtcontaining JSON-like content.gmt
Custom resources are parsed by the application and used in the same enrichment, distance calculation, and visualization workflow as the built-in resources.
The application can export an output file named:
analysis_results.tsv
The exported result file includes fields such as:
- Gene set name
- p-value
- q-value or adjusted p-value
- Enrichment direction
- Gene set size
- Matched genes
GSPlot follows this general workflow:
- Parse the user-provided gene input.
- Load the selected gene set resource.
- Filter gene sets based on the minimum matched-gene requirement.
- Run enrichment analysis according to the selected input mode.
- Adjust or report statistical significance values.
- Keep gene sets that pass the selected threshold.
- Compute pairwise gene-set similarity or distance.
- Apply dimensionality reduction using UMAP, t-SNE, or Isomap.
- Render the significant gene sets as an interactive 2D graph.
- Support filtering, point selection, clustering, labeling, and export.
GSPlot supports multiple input modes because users may have different types of gene-level data.
Ranked mode accepts an ordered gene list. The order of the genes is used to test whether members of a gene set are concentrated toward the selected side of the list.
Depending on the selected hypothesis, GSPlot can test for enrichment toward the positive side, negative side, or both sides of the ranked input.
Thresholded mode accepts significant and insignificant/background gene lists. GSPlot uses Fisher's exact test to evaluate whether each gene set contains more relevant genes than expected.
This mode is useful when the user already has a selected list of significant genes from an earlier analysis.
Scored mode accepts genes with numerical scores. GSPlot uses a preranked enrichment workflow through GSEApy to evaluate gene set enrichment based on the score ordering.
This mode is useful when each gene has a continuous value, such as a differential expression score, test statistic, or other ranking metric.
After enrichment analysis, GSPlot compares significant gene sets based on their member overlap. Gene sets with more shared genes are treated as more similar and are placed closer together in the final visualization.
Supported or planned similarity/distance options may include:
- Jaccard distance
- Overlap coefficient
- Weighted variants where applicable
The resulting distance matrix is used as input for dimensionality reduction.
GSPlot supports several dimensionality-reduction methods for placing gene sets in a 2D plot:
- UMAP
- t-SNE
- Isomap
The final coordinates are used only for visualization. Similar or overlapping gene sets should appear near each other, but the exact layout can vary depending on the selected method, parameters, software versions, and random seed behavior.
Le TC, Espinoza CG, Courtois J, Babur Ö. GSPlot: Visualizing gene set similarity in enrichment results. bioRxiv. 2026 Jun 16:2026-06.
To improve reproducibility:
- Use the pinned dependency versions in
requirements.txt. - Record the selected input mode.
- Record the selected directional hypothesis.
- Record the selected gene set resource and version.
- Record the selected gene set collections.
- Record the p-value or q-value threshold.
- Record the minimum matched-gene requirement.
- Record the selected similarity/distance metric.
- Record the selected dimensionality-reduction method and relevant parameters.
- Use the same code release, commit hash, or version tag when reproducing published results.
Dimensionality-reduction layouts may differ across software versions, random seeds, and computing environments, even when the enrichment results are the same.
Do not commit sensitive or private files to the repository.
Examples of files and information that should not be committed:
.envfiles- API keys
- access tokens
- Django secret keys
- private datasets
- restricted licensed datasets
- local database files
- user-uploaded private data
Use .env.example or documentation to show required environment variable names without exposing real values.
The live deployment is available at:
https://gsplot.cs.umb.edu
Deployment settings may differ from local development settings. Production deployments should use:
- a secure Django secret key
DEBUG=False- allowed host configuration
- a production web server setup
- protected environment variables
- appropriate timeout and upload-size settings for larger gene set analyses
For questions, bugs, or feature requests, please open an issue in this GitHub repository.
Project team:
Tien Le, developer
Ozgun Babur, mentor
Network Biology Lab:
https://sites.google.com/view/umb-network-biology
Live app/demo:
https://gsplot.cs.umb.edu
GitHub repository:
https://github.com/PathwayAndDataAnalysis/gsplot
Network Biology Lab:
https://sites.google.com/view/umb-network-biology
Paper, preprint, or documentation links can be added here when publicly available.