ProfileBuilder
Multi-threaded orchestrator — builds a DatasetProfile from a CSV file.
// include/zedda/profile_builder.hpp
namespace zedda {
class ProfileBuilder {
public:
using ProgressCallback = std::function<void(int64_t rows_done)>;
explicit ProfileBuilder(const std::string& path, StreamReaderConfig config = {});
void set_progress(ProgressCallback cb);
DatasetProfile build(bool is_sampled = false,
int64_t sample_size = 1000000,
bool correlate = false);
};
} // namespace zedda
Purpose
Multi-threaded orchestrator. Builds a DatasetProfile from a CSV file by:
- Constructing a
CsvStreamReaderfor the file. - Allocating per-column
ColumnAccumulatorinstances. - Splitting the file into chunks and dispatching them to a thread pool (vendored
BS::thread_pool). - Merging per-chunk accumulators into per-column accumulators via
ColumnAccumulator::merge(). - Computing pairwise correlations on numeric columns (unless skipped).
- Finalising each column’s statistics via
ColumnAccumulator::finalize(). - Returning the assembled
DatasetProfile.
Methods
set_progress(cb)
Installs a progress callback invoked after each chunk with the number of rows processed so far. Called from worker threads — your callback must be thread-safe.
build(is_sampled, sample_size, correlate)
Runs the full pipeline. Parameters:
| Parameter | Default | Meaning |
|---|---|---|
is_sampled |
false |
If true, marks the resulting profile as sampled |
sample_size |
1000000 |
Maximum rows to read when is_sampled is true |
correlate |
false |
If true, force O(N²) correlation even when numeric cols > 50 |
Returns a DatasetProfile.
Correlation skipping (PERF-1)
By default (correlate=false), the builder skips the O(N²) correlation matrix when the dataset has more than 50 numeric columns. The resulting DatasetProfile has correlation_skipped = true and an empty correlations vector.
Pass correlate=true to force correlation. The Arrow profiler has a separate hard cap at MAX_CORR_COLS = 1000 (SEC-C01) — above 1,000 numeric columns, correlation is always skipped.
Python entry point
The Python scan() / profile() entry point is fasteda_core.profile(path, show_progress=True, is_sampled=False, sample_size=1000000, correlate=False), which constructs a ProfileBuilder and calls build(). The GIL is released for the duration of the scan.
See also
- C++ API: CsvStreamReader — what
ProfileBuilderreads from. - C++ API: ColumnAccumulator — what
ProfileBuilderfeeds. - C++ API: CorrelationEngine — how correlation is computed.
- Architecture — the threading model.