Visual exploration and clustering tool for image embeddings. Users can either bring pre-calculated embeddings to explore, or use the interface to embed their images and then explore those embeddings.
| Embed & Explore | Precalculated Embedding Exploration |
![]() |
![]() |
![]() |
![]() |
![]() |
Embed & Explore - Embed images using pretrained models (CLIP, BioCLIP), cluster with K-Means, visualize with PCA/t-SNE/UMAP, and repartition images by cluster.
Precalculated Embeddings - Load parquet files (or directories of parquets) with precomputed embeddings, apply dynamic cascading filters, and explore clusters with taxonomy tree navigation. See Data Format for the expected schema and Backend Pipeline for how embeddings flow through clustering and visualization.
Requires Python 3.10–3.13. We recommend using uv for package management.
git clone https://github.com/Imageomics/emb-explorer.git
cd emb-explorer
# Create venv with Python 3.12 (3.14+ not yet supported by GPU dependencies)
uv venv --python 3.12 && source .venv/bin/activate
uv pip install -e .A GPU is not required — everything works on CPU out of the box. But if you have an NVIDIA GPU with CUDA, clustering and dimensionality reduction (KMeans, t-SNE, UMAP) will be significantly faster via cuML.
# CUDA 12.x
uv pip install -e ".[gpu-cu12]"
# CUDA 13.x
uv pip install -e ".[gpu-cu13]"The app auto-detects GPU availability at runtime and falls back to CPU if anything goes wrong — no configuration needed. The CPU sklearn path is auto-accelerated by scikit-learn-intelex1. You can also manually select backends (cuML, sklearn) in the sidebar.
To get reproducible projections and clusters, enable Use fixed seed in the sidebar and pin the backend instead of auto: the GPU backend is cuML, the CPU backend is sklearn (auto-accelerated by scikit-learn-intelex on x86 CPUs). With a seed and a pinned backend, results are identical across app restarts on both backends, with one exception: cuML t-SNE, which never reproduces exactly.
- PCA is a deterministic decomposition, with no stochastic optimization involved. cuML PCA uses a full eigendecomposition and always returns the same result, seed or no seed. sklearn can auto-select a randomized SVD solver, so the app passes the seed to make it reproducible.
- UMAP and KMeans reproduce exactly on both backends when a seed is set. (Seeded UMAP trades some speed for determinism.)
- t-SNE reproduces on
sklearnwhen a seed is set. cuML t-SNE does not reproduce, by a deliberate trade-off:cuML ≥ 26.08made its default FFT solver deterministic under a fixed seed, but the FFT solver collapses to a degenerate ~1D line on highly homogeneous embeddings (#40), therefore our application pins cuML t-SNE to theexactsolver, which is collapse-free but remains non-deterministic (GPU floating-point accumulation order varies between runs, and t-SNE's optimization amplifies the difference). Selectsklearnwhen t-SNE results need to be reproducible. autochooses a backend from data size and hardware, so the same seed can run different algorithms on different machines. Exact coordinates may also differ across library versions and hardware; a seed guarantees repeatability within one environment, not across environments.
# Embed & Explore - Interactive image embedding and clustering
streamlit run apps/embed_explore/app.py
# Precalculated Embeddings - Explore precomputed embeddings from parquet
streamlit run apps/precalculated/app.pyemb-embed-explore # Launch Embed & Explore app
emb-precalculated # Launch Precalculated Embeddings app
list-models # List available embedding modelsAn example dataset (data/example_1k.parquet) is provided with BioCLIP 2 embeddings for testing. Please see the data README for more information about this sample set.
# On compute node
streamlit run apps/precalculated/app.py --server.port 8501
# On local machine (port forwarding)
ssh -N -L 8501:<COMPUTE_NODE>:8501 <USER>@<LOGIN_NODE>
# Access at http://localhost:8501Footnotes
-
sklearn-intelexis powered by the oneDAL library that provides accelerations on x86_64 Linux and Windows machines, and silently fall back to vanillasklearnon unsupported architectures like Apple Silicon and ARM Linux. The package is under the UXL Foundation (a Linux Foundation project) so cross-vendor support is a stated goal. ↩




