1
Load a dataset
Drop a CSV here, or click to browse
Numeric and categorical columns are scaled and encoded automatically
Or try:
2
Choose how many clusters
Columns used identifiers and constants are excluded automatically
Silhouette by k higher is better separated — the peak is a good starting point
Clustering, and running the shuffled-null comparison…
3
Are these clusters real?
Projection first two principal components — a 2-D shadow of a higher-dimensional space
Every algorithm, same data
4
What makes each cluster different
Not a table of averages — the features that sit furthest from the overall population, measured in standard deviations. That is what actually characterises a segment.