C++ API

Embed the BlazeRules C++ core directly — construct the engine, load rules, evaluate batches, and read decisions without Python in the hot path.

The Python module blazerules is a thin binding over the C++20 core. Native applications can link the core directly to produce a single binary, integrate with an existing C++ service, or minimize per-batch binding overhead. This reference maps public headers and engine methods to their Python equivalents.

📘

Same engine, two front doors

The C++ core and Python module use the same compiler, kernels, schema inference, and decision logic. Evaluation semantics are shared across both interfaces; call syntax and several enum names differ.

Headers

All core headers live under include/blazerules/:

HeaderAPI surface
engine.hRuleEngine, EngineConfig, BacktestConfig, RuleFileFormat
schema.hColumnType, FieldSpec, BlazeRulesSchema
batch_result.hBatchResult, BacktestReport, timing and per-rule structures
conflict.hConflictReport — returned by load_rules
rule_spec.hRule, action, and operator types shared across the API

engine.h includes <arrow/api.h>, so the Apache Arrow development headers must be available on the include path.

The shape of the engine

A RuleEngine compiles a YAML ruleset into immutable execution plans and evaluates record batches against that plan. Each call returns a BatchResult with the materialized fields selected by output_detail. Schema may be supplied explicitly or inferred from the first evaluated batch.

#include "blazerules/engine.h"

EngineConfig cfg;
cfg.output_detail = EngineConfig::OUTPUT_DECISIONS;   // see enum note below

RuleEngine engine(cfg);                               // non-copyable; own it / pass by reference
ConflictReport report = engine.load_rules("rules.yaml");
// inspect `report` before serving — load is strict (see below)

std::string ndjson = /* one JSON object per line */;
BatchResult r = engine.evaluate_ndjson(ndjson);

for (size_t i = 0; i < r.decisions.size(); ++i) {
    // r.decisions[i], r.scores[i], r.winning_rule_ids[i] ...
}

Constructors

  • explicit RuleEngine(EngineConfig config = {}) — default or configured engine; schema inferred on first batch.
  • RuleEngine(std::vector<FieldSpec> fields, EngineConfig config = {}) — bind a schema up front.
  • static BlazeRulesResult<std::unique_ptr<RuleEngine>> RuleEngine::create(BlazeRulesSchema schema, EngineConfig config) — factory returning a result wrapper instead of throwing.

RuleEngine is non-copyable (the copy constructor and copy assignment are deleted). Move it or pass it by reference; never copy it into a container by value.

Loading rules

ConflictReport load_rules(const std::string& rules_path);
ConflictReport load_rules_from_string(const std::string& rules_yaml_or_json,
                                      RuleFileFormat format = RuleFileFormat::YAML); // enum class RuleFileFormat { YAML, JSON }
ConflictReport reload_rules_now(const std::string& rules_path);
ConflictReport analyze_conflicts(const std::string& rules_path);
std::string    active_rule_set_version() const;
🚧

load_rules returns a report, not void

Activation is strict: malformed YAML, duplicate rule IDs, and missing lookup references fail before activation. load_rules returns a ConflictReport describing detected overlaps and conflicts.

Evaluating batches

BatchResult evaluate_ndjson(std::string_view ndjson_bytes);
BatchResult evaluate_ndjson_padded(std::string_view ndjson_bytes);   // input already simdjson-padded
BatchResult evaluate_ndjson_file(const std::string& path);           // mmap, zero-copy replay
BatchResult evaluate_json_array(std::string_view json_array_bytes);  // top-level JSON array
BatchResult evaluate_json_array_padded(std::string_view json_array_bytes);
BatchResult evaluate_record_batch(const std::shared_ptr<arrow::RecordBatch>& batch);
BatchResult evaluate_batch(const std::shared_ptr<arrow::RecordBatch>& batch); // inline alias for evaluate_record_batch
BatchResult evaluate_messages(const std::vector<std::string>& messages);
BatchResult evaluate_message_views(const std::vector<std::string_view>& messages);

Prefer evaluate_batch when upstream data is already typed Arrow,
evaluate_ndjson for newline-delimited JSON streams, and
evaluate_json_array when the input is already one top-level JSON array. The
JSON-array entry point traverses the array directly instead of serializing it
back into NDJSON. Each evaluator has an _into(..., BatchResult& out) variant
that reuses an existing result object to avoid per-batch allocation in tight
loops.

Advanced: sharding and partition affinity

For window-heavy streaming workloads, entity affinity can be preserved across shards. create_shards(int) returns owned per-shard engines, and the evaluate_partition(int partition_id, ...) overloads evaluate one partition's records. Sharding is appropriate when profiling identifies window-state contention or a single-engine throughput limit.

Models, hot reload, schema, backtest

void register_model(const std::string& name, const std::string& path);  // ONNX
int  num_models() const;

void enable_hot_reload(const std::string& rules_file_path,
                       std::chrono::seconds poll_interval = std::chrono::seconds(5));
void stop_hot_reload();
HotReloadStatus hot_reload_status() const;

const BlazeRulesSchema& schema() const;
bool schema_bound() const;
SchemaState schema_state() const;   // enum class SchemaState { UNBOUND, INFERRED_BOUND, USER_BOUND }

BacktestReport backtest(const BacktestConfig& config);
🚧

ONNX is build-gated

model_score rules and register_model(...) require an ONNX-enabled build (BLAZERULES_ENABLE_ONNX, default ON). In a build with ONNX off, model_score rules are rejected at compile time and register_model throws. See Backtesting a Candidate for the backtest(...) workflow.

The methods above are the load → evaluate → read surface most embedders need. engine.h also exposes reset_window_state(), num_window_channels(), in-process metrics (enable_metrics(), reset_metrics(), metrics_counters(), metrics_gauges(), metrics_histograms()), stats(), and partition-affine evaluation helpers.

C++ vs Python, side by side

import blazerules

config = blazerules.EngineConfig()
config.output_detail = blazerules.OutputDetail.DECISIONS

engine = blazerules.RuleEngine(config)
engine.load_rules("rules.yaml")

result = engine.evaluate_ndjson(ndjson_bytes)
for decision in result.decisions:
    ...

Enum names differ between the two front doors

The EngineConfig modes are nested unscoped enums and use EngineConfig::<NAME>. Constant names differ from the Python enums:

SettingPythonC++
Output detailOutputDetail.DECISIONSEngineConfig::OUTPUT_DECISIONS
Output detailOutputDetail.BITMASKSEngineConfig::OUTPUT_BITMASKS
Ingest errorsIngestErrorMode.SKIP_AND_COUNTEngineConfig::SKIP_AND_COUNT
Type mismatchTypeMismatchMode.NULL_ON_TYPE_ERROREngineConfig::NULL_ON_TYPE_ERROR
📘

The C++ default is OUTPUT_BITMASKS

In C++, EngineConfig::output_detail defaults to OUTPUT_BITMASKS. OUTPUT_DECISIONS avoids per-rule mask materialization when only routing results are required.

Reading a BatchResult in C++

The data is the same as the Python result, but a few members that are methods in Python are plain fields in C++:

DataPythonC++
Records / matchesn_records, n_matchedn_records, n_matched (fields)
Decisionsdecisions, scores, winning_rule_idssame field names
Grouped indicesgrouped_decision_indices() (method)grouped_decision_indices (member field)
Per-rule countsmatch_countsrule_match_counts (member field)
Matched rowsmatched_indicesmatched_record_indices (member field)
Timingtiming_mstiming_ms() (methodunordered_map<string,double>)
BatchResult r = engine.evaluate_ndjson(ndjson);

int matched = r.n_matched;
const auto& approve_rows = r.grouped_decision_indices["APPROVE"];   // member access
auto timings = r.timing_ms();                                       // method in C++ too
double total_ms = timings["total"];

Zero-copy buffer helpers (rule_bitmask_buffer(id), matched_indices_buffer(), decision_codes_buffer(), grouped_indices_buffer(decision)) expose the underlying memory as arrow::Buffer for transfer to another Arrow consumer without copying.

Linking against the core

The core CMake target is blazerules_core:

# either find an installed package...
# find_package(blazerules CONFIG REQUIRED)
# ...or add the repo as a subdirectory
add_subdirectory(third_party/blazerules)

add_executable(app main.cpp)
target_link_libraries(app PRIVATE blazerules_core)

The install export namespace is blazerules::, so installed consumers can link blazerules::blazerules_core when using the generated CMake package. In-tree builds link the target as blazerules_core.

Next steps


Did this page help you?