Practical tactics to build reliable AI apps — Dmitry Kuchin, Multinear
AI Engineer World's Fair 2025 · 14:55
AI consulting and evaluation software
Multinear provides AI strategy, implementation and optimization services alongside an AI evaluation platform for engineers and product managers. Its client work includes support chatbots with expert-reviewed guardrails, document question-answering for employee onboarding, and text-to-SQL dashboards for nontechnical teams. The software lets teams benchmark application changes and detect regressions across prompts, models, data and business logic.
Founded by technology leaders with backgrounds at Meta, eBay, ZoomInfo and Northwestern Mutual, Multinear emphasizes hands-on engineering. Its evaluation approach uses business-oriented pass/fail tests to make application behavior measurable. Teams can assess outputs through strict comparisons, LLM judges or human evaluations, then compare experiment runs to identify where changes improve results or break previously successful cases.
The services cover work from proof of concept through production implementation, plus improvements to existing systems’ accuracy, speed and cost efficiency. The evaluation software is distributed as an MIT-licensed Python package that runs locally through a browser interface or command line and stores experiment results in SQLite. This gives teams a way to keep evaluation configuration and results alongside their application code.
AI Engineer World's Fair 2025 · 14:55
Affiliations reflect their AIE appearances, not necessarily current employment.
Kuchin describes how iterative tests support model and architecture comparisons, including LLM-as-a-judge workflows and mock-database testing for Text-to-SQL.
Affiliations reflect each recorded session, not necessarily current employment.