Charles Dickens researches how explicit knowledge and expert judgment can improve machine learning. He co-developed NeuPSL, which combines neural predictions with symbolic rules, and has co-authored work on graph compression, specialized agents, and the reliability of expert evaluation. At the AI Engineer World’s Fair 2026, he was a Senior Applied Research Scientist at Snorkel AI.
From electrical engineering to structured learning
Dickens earned a bachelor’s degree in electrical and computer engineering, with a mathematics minor, at the University of Hawai‘i at Mānoa in 2019. His early research included classifying electrical disturbances for the Open Power Quality project. He subsequently joined the LINQS research group at UC Santa Cruz, working with Lise Getoor on statistical relational learning. During an applied-scientist internship at Amazon in 2022, he studied how to train graph neural networks on large datasets. His early career connected mathematical modeling with practical machine-learning software.
At Santa Cruz, Dickens developed methods for predictions that depend on one another. A recommendation, a graph label, or a demand forecast can change when related information arrives; recomputing an entire model each time can be expensive. His co-authored online collective inference research, published in 2021, addressed evolving graphical models through selective updates, with theoretical guarantees on stability and regret. Experiments covered recommendation, trust prediction, and demand forecasting.
His doctoral work extended this focus to the relationship between learned perception and explicit knowledge. His August 2024 dissertation developed a mathematical framework for neural-symbolic systems, modeling patterns, learning algorithms, and a practical implementation in NeuPSL. The framework organized hybrid architectures by how their neural and symbolic components interact and how the combined system learns.
Rules, graph structure, and disciplined tool use
NeuPSL and joint reasoning: Dickens co-developed Neural Probabilistic Soft Logic with Connor Pryor and collaborators. Neural networks supply perceptual predictions; symbolic rules express relationships that the combined output should respect. An energy function measures their compatibility, and inference searches for a low-energy solution. In one experiment, handwritten digits recur across several addition problems. Even without individual digit labels, the known sums constrain their possible identities across equations: evidence from one problem can resolve ambiguity in another. The experimental repository also includes visual Sudoku and citation-network classification.
Graph compression that preserves useful computation: Dickens is the first-listed author of ConvMatch, developed through his Amazon research. The method summarizes a graph while aiming to preserve graph-convolution outputs—the results of combining information from connected nodes. It selects node merges according to their effect on that computation. Across the paper’s node-classification experiments, models trained on graphs reduced to 1% of their original size retained up to 95% of the original prediction performance. The result is specific to those experiments, rather than a general compression guarantee.
Specialization through tool discipline: In work with Snorkel and UC Berkeley’s Sky Computing Lab, Dickens described training a Qwen3 4B financial agent with reinforcement learning for under $500. On the reported financial-question benchmark, it reached about 60% accuracy versus 51% for its 235B sibling. The team identified failures such as hallucinated table schemas, excessive context accumulation, and repeated unsuccessful strategies. Simpler training data and a binary pass/fail reward worked best; the improvement centered on disciplined tool use. The learned skills transferred to harder questions involving multiple tables without reducing general tool-use performance in the reported tests. The team released the rLLM-FinQA-4B model, trained using the open-source rLLM framework.
Testing professional work—and the judgments used to score it
Dickens’s enterprise-evaluation research examines the information gathering and judgment that professional tasks require. At Snorkel, he co-authored UNDERWRITE, an insurance-underwriting benchmark built with domain experts. Its tasks include business-specific knowledge, noisy tool interfaces, and simulated users who require follow-up questions. These features make gathering the right information part of the test. The team found that accuracy and efficiency did not move together, and that models could invent domain knowledge despite having tools available to retrieve it.
He also studies whether the scoring process captures the quality that matters to practitioners:
Measuring rubric quality: In writing co-authored with Chris Glaze, Dickens argues that scoring rubrics should themselves be measured and improved. Graders can agree while applying criteria that miss stakeholders’ objectives; useful criteria can also fail if people cannot apply them consistently. His co-authored RIFT research organizes rubric defects into reliability, content-validity, and consequential-validity failures.
Comparing ways to elicit expertise:JudgmentBench, which Dickens co-authored, tests whether rubric scoring is the best way to obtain expert assessments. In its initial comparison of legal outputs, practicing attorneys’ pairwise judgments recovered the constructed quality ordering better than rubric scores and took less than half the annotation time. That result gives a concrete reason to evaluate the method used to collect expert judgment, alongside the AI outputs being judged.
Charles Dickens explains how verified financial data and reinforcement learning helped a 4-billion-parameter agent outperform its 235-billion-parameter sibling—and why disciplined tool use mattered more than harder training examples or elaborate rewards.
On 290 held-out financial questions, the trained 4B model reached almost 60% accuracy versus 51% for its 235B sibling. The advantage is specific to this simulated financial task.
The reported training run cost about $420 for compute and $40 for judging. A binary final-answer reward outperformed the more elaborate reward designs tested.
Single-table training produced the greatest lift among the tested data mixtures and transferred to questions involving two to five tables, while general tool-calling performance held up.