Formula Evaluation and Custom Evaluators¶
An Evaluator is the shared evaluation protocol used by formula-evaluation tools in SRHarness. It encapsulates how data is split, how expression parameters are fitted, and which numeric metrics are used to evaluate the fitted expression. Tools such as evaluate_formula, submit_formula, polynomial fitting, and symbolic-search backends can therefore share one evaluation policy.
The active instance is stored in AgentContext.evaluator. SRHarness provides DefaultEvaluator for ordinary aligned data and GraphEvaluator for graph and hypergraph data. A user-defined Evaluator can inherit DefaultEvaluator and override only the behavior required by the task.
The role of SRHarness Engine
Evaluators use the SRHarness Engine Expression representation together with its parsing, parameter-fitting, and numerical-evaluation capabilities. The Engine does not decide data splits, candidate eligibility, metric definitions, or ranking policy; those belong to Evaluators and formula-evaluation tools. Keeping mathematical-expression infrastructure in the Engine lets tools and custom Evaluators share one mathematical semantics.
Evaluation path¶
A formula tool first parses model-provided text into an SRHarness Engine Expression. BaseTool.evaluate() then determines whether the formula is a formal candidate:
- when the left-hand expression is exactly
context.targetand the right-hand side does not depend on the target, it is a candidate and usesfit_candidate()andevaluate_candidate(); - other equalities remain available for relationship exploration, but use the general
fit()andevaluate()path and do not trigger candidate-only behavior.
flowchart TD
A[Formula-evaluation tool] --> B[Parse f and y]
B --> C[Obtain train and validation contexts]
C --> D{Formal candidate?}
D -- Yes --> E[fit_candidate]
D -- No --> F[fit]
E --> G[evaluate_candidate]
F --> H[evaluate]
G --> I[Train and validation metrics]
H --> I
I --> J[ToolCallResult and candidate ranking]
The following pseudocode summarizes how formula tools use the five entry points:
def evaluate_formula(f, y, context, evaluator):
split = evaluator.split(context)
train_context = split["train"]
validation_context = split["validation"]
is_candidate = (
y.to_str() == context.target
and context.target not in f.variables
)
fitted = (
evaluator.fit_candidate(f, train_context)
if is_candidate
else evaluator.fit(f, y, train_context)
)
evaluate = (
evaluator.evaluate_candidate
if is_candidate
else lambda expression, split_context: evaluator.evaluate(
expression, y, split_context
)
)
return {
"train": evaluate(fitted, train_context),
"validation": evaluate(fitted, validation_context),
}
When fitting is requested, parameters are determined only on the training split. The same fitted expression is then evaluated on the training and validation splits. On first access, AgentContext.train_split or AgentContext.validation_split calls evaluator.split(context) and caches both context views until the data changes or the cache is explicitly invalidated.
Residual diagnostics, candidate eligibility, result formatting, and error wrapping remain responsibilities of BaseTool; individual Evaluators do not need to reimplement them.
The DefaultEvaluator contract¶
DefaultEvaluator defines five class methods. A custom Evaluator may override any subset:
| Method | Input context | Responsibility |
|---|---|---|
split(context) |
Complete, unsplit context | Return exactly {"train": AgentContext, "validation": AgentContext}. |
fit(f, y, context) |
Training context | Fit parameters in f to approximate expression y, then return an Expression with bound parameters. |
evaluate(f, y, context) |
Training or validation context | Evaluate a general equality y = f and return numeric metrics. |
fit_candidate(f, context) |
Training context | Fit a formal candidate; by default this delegates to fit(f, Symbol(context.target), context). |
evaluate_candidate(f, context) |
Training or validation context | Evaluate a formal candidate; by default this delegates to evaluate(f, Symbol(context.target), context). |
This layering preserves the Agent's ability to explore general equalities while giving formal candidates a task-specific extension point. For example, a dynamics Evaluator can integrate dx_dt = f(x, t) and calculate trajectory error only in evaluate_candidate() without trying to integrate every implicit equality used for analysis.
An Evaluator must obey these result contracts:
fit()andfit_candidate()returnsr_harness_engine.Expressionobjects with no unbound parameters at evaluation time;evaluate()andevaluate_candidate()returndict[str, float | int]; booleans, strings, arrays, and nested values are not valid metrics;- metric names enter train and validation results unchanged, and any metric can be selected as the candidate-ranking key;
complexity, like every other metric, is returned by the Evaluator rather than injected separately by the runtime.
Built-in Evaluators¶
DefaultEvaluator¶
DefaultEvaluator assumes sample arrays in context.data are aligned along their first dimension. It supports random and OOD splitting, uses SRHarness Engine to fit expression parameters and evaluate predictions, and returns these metrics by default:
GraphEvaluator¶
GraphEvaluator derives from DefaultEvaluator and handles data composed of (..., N), (..., E), and (..., H) arrays plus graph or hypergraph relation tables. It passes context.num_nodes during fitting and evaluation and slices only sample-aligned dimensions, leaving shared edge and hyperedge tables intact.
Evaluator utilities¶
Reusable metric, splitting, and dynamics infrastructure is available from:
It includes:
calc_MSE(),calc_RMSE(),calc_MAE(),calc_MAPE(), andcalc_R2();- AIC, BIC, Pearson/Spearman correlation, and expression complexity;
- random, OOD, and chronological data splitting;
integrate_ODE()andcalc_trajectory_rollout_RMSE().
Keeping these reusable operations in utils makes custom Evaluators concise without expanding the Evaluator contract into a general infrastructure layer.
Custom Evaluators¶
The safest extension pattern is to inherit DefaultEvaluator and override the narrowest entry point that needs to change. This Evaluator preserves every default behavior and adds trajectory rollout RMSE only for formal dynamics candidates:
from . import utils
from .default_evaluator import DefaultEvaluator
class TrajectoryRolloutEvaluator(DefaultEvaluator):
@classmethod
def evaluate_candidate(cls, f, context):
metrics = super().evaluate_candidate(f, context)
metrics["rollout_rmse"] = utils.calc_trajectory_rollout_RMSE(
f,
context,
time="t",
state="x",
)
return metrics
Override fit_candidate() as well when the fitting objective itself should include rollout behavior. Override fit() or evaluate() only when both general equalities and formal candidates require a different base policy.
A custom source file must define exactly one DefaultEvaluator subclass. load_custom_evaluator() loads and inspects the source in an isolated process, returns a same-named protocol proxy in the main process, and retains source/file provenance. Subsequent split, fit, and evaluate calls also run through the operating-system sandbox; custom source is never executed directly inside the Web service process. The WebUI can save, load, and test scripts in context.evaluator/, while the Evaluator Construction Agent can create or repair them; see SRHarness WebUI.
See the API Reference for complete signatures and SRHarness Engine for expression syntax, parameters, and graph evaluation rules.