Your NILM models are starving for data.
Public datasets aren't enough anymore.
Real multi-site, multi-sector industrial energy datasets. Machine-level ground truth, high sampling rate, three-phase measurements. The scientific baseline for load disaggregation and predictive maintenance.
# Load a dataset in three lines — NILMTK-compatible out of the box.
from nilmtk import DataSet
ds = DataSet("watt_industrial_v2.h5")
ds.set_window(start="2025-01-01")
elec = ds.buildings[3].elec
# > MeterGroup(meters=42, sample_period=0.02s, sectors=['mechanical'])// The gap
Industrial NILM data is scarce, short, and mono-context.
HIPE, IMDELD, SIDED — you can count public industrial datasets on one hand. Three months of data, a single site, a single sector. Models trained on them don't generalize. Collecting proprietary data costs months and tens of thousands of euros in hardware, installation and labeling.
// The dataset
Engineered for load disaggregation research.
Every spec is measurable. No marketing padding.
V, I, P, Q, S, harmonics, THD, power factor per phase.
Submetering per single machine, not per circuit.
Mechanical, food & beverage, plastics, packaging, HVAC.
Seasonal variation and full production cycles captured.
On/off transitions, state changes, anomalies, standby.
HDF5 + CSV. Documented metadata schema. Ready for your pipeline.
Anonymized. Clear license for commercial model training.
Residential appliance-level data to complement industrial.
// Comparison
Us vs. public datasets
Feature parity with HIPE, IMDELD, SIDED — plus everything they lack.
| Ours | HIPE | IMDELD | SIDED | |
|---|---|---|---|---|
| Duration | ≥ 12 months | 3 months | 6 months | 4 months |
| Sites | 12 | 1 | 1 | 1 |
| Machines | 180+ | 10 | 8 | 11 |
| Sectors | 5 | 1 | 1 | 1 |
| Sampling rate | 1 Hz | 5 kHz | 1 Hz | 1 Hz |
| Machine-level labels | Yes | Partial | Yes | Partial |
| Event labeling | Yes, > 2M events | No | Limited | No |
| Commercial license | Yes | No | No | No |
| Technical support | Included | None | None | None |
// Use cases
Built for the workloads NILM teams actually run.
Train FHMM, seq2seq, CNN, transformer architectures on realistic multi-machine mixtures.
Reproducible baselines with known ground truth. Compare against SOTA on the same conditions.
Electrical signatures across full production cycles for anomaly detection and failure precursors.
Cross-site and cross-sector data for domain adaptation to unseen machinery families.
// Methodology
Every measurement is documented, reviewed, and traceable.
We ship a datasheet with each dataset. QA is not optional.
Class 0.5 revenue-grade meters, PT/CT chain calibrated, timestamped via PTP/NTP with drift monitoring.
Dual-review labeling with operator interviews. Event schema versioned in Git.
Automated coverage checks, gap analysis, energy-balance reconciliation, anomaly flagging.
On-demand acquisition on your machinery, your protocol, your license terms.
// Pricing
Three tiers.
Sample and academic tiers are self-serve. Commercial is a real conversation.
For evaluation. Enough data to validate your pipeline before you commit.
- ✓1 site · 5 machines · 7 days
- ✓HDF5 + CSV export
- ✓NILMTK loader
- ✓Community support
For universities and public research labs. Full dataset access, non-commercial.
- ✓Full multi-site coverage
- ✓Machine-level ground truth
- ✓Event labels + metadata
- ✓Email support · 48h SLA
- ✓Citation guidelines
For product teams shipping NILM in production. Includes custom acquisition on request.
- ✓Everything in Research
- ✓Commercial training license
- ✓Custom sites & sectors
- ✓Dedicated technical contact
- ✓Update stream (quarterly)
// Get in touch
Request a sample or a technical call.
We reply within one business day. No sales funnel, you talk to engineers.
30 min with an engineer. Pick a slot that works for you.
Book a technical call