acceleratedlogicai.com/logicmark
Objective model benchmarks

LogicMark

Run fixed Math and Coding suites with your equipped models, then compare your personal average scores and attempt counts.

  • ›Fixed versioned Math and mixed-language Coding suites.
  • ›A personal leaderboard with average scores and attempt counts.
  • ›Deterministic grading with first-party isolated runtimes.

Choose a model and run a suite

LogicMark uses models equipped in Model Organizer and their configured providers. Math and Coding each have one run button. Every attempt uses the same versioned task set. Tool-assisted mode starts a fresh conversation and workspace per task; Quick mode sends the entire suite and starter files in one request and grades one response. Benchmark selection does not switch the model active in your chat workspace.

Deterministic checks

The Math suite has twelve exact-answer problems and an optional Python tool in Tool-assisted mode. The Coding suite has ten practical starter-repository tasks across Python and JavaScript. In Tool-assisted mode, models can read and write those files; in Quick mode they return their final files together. An isolated checker executes submissions after they are complete and compares outputs to expected results. Grading uses code, never an AI judge.

Your personal leaderboard

The leaderboard groups models by test, mode, and suite version, showing their mean score across completed attempts and their attempt count. You can repeat runs without a LogicMark attempt limit. Provider quotas and runtime limits still apply, and provider or evaluator outages do not record a score. These small task suites measure specific behavior and do not establish a general intelligence ranking.

Scores and backups

Only aggregate scores, model labels and attempt counts are stored in your browser. Model responses, submitted files and tool transcripts are discarded after a run. Scores are included in full database exports and optional manual Drive snapshots. Clear Results removes current local scores and, when Drive is connected, removes them from the latest app-managed snapshot while preserving its other data. Downloaded backups remain under your control.

Frequently asked questions

How are scores calculated?

Each task contributes equally to its suite score. Math answers are compared exactly; coding tasks receive credit for passing deterministic checks. The leaderboard displays the mean suite score over completed attempts.

Which models can I test?

Models equipped in Model Organizer, including supported local and custom provider entries. Requests use your chosen provider configuration.

Can coding models run tests themselves?

Tool-assisted mode allows reading and writing files. Quick mode supplies all starter files and accepts final files in one response. Execution belongs to the isolated checker after submission, and hidden expected results are not provided to the model.

Can I repeat or clear my scores?

Yes. You can run another attempt after a suite finishes, stop a running suite, or clear saved score totals. Stopped and infrastructure-failed suites do not affect your average.

Choose a focused workspace, then return to the catalog to compare more tools.