Gabriel Chua is OpenAI’s Developer Experience Engineer for APAC, helping developers build, ship, and scale with Codex. His work spans developer tools, open-source AI applications, and responsible-AI research shaped by Singapore’s languages, public services, and users.
From public-sector analytics to developer enablement
Chua previously worked on healthcare finance as a policy analyst at Singapore’s Ministry of Health before becoming a data scientist at GovTech. His public-sector work grew to include applied responsible-AI research and machine-learning and language-model systems to combat online scams. These applications required attention to the people using a system, the language they used, and the consequences of an incorrect response.
Alongside that technical work, Chua helped build a community for people developing AI applications. He co-started AI Wednesdays with collaborators, connecting that effort with Lorong AI. Community building forms a meaningful part of his transition from public-sector AI to developer enablement: his work includes both building systems and helping other practitioners develop their own.
Safety designed for particular users
Chua’s collaborative research examines problems that broad safety categories can obscure. A request can be harmless yet outside an application’s purpose; a moderation system can perform differently across local languages; and an answer appropriate for adults can be unsuitable for children. His work makes those distinctions concrete through guardrails, benchmarks, and testing methods.
Application-specific guardrails: With Shing Yee Chan and Shaun Khoo, Chua co-authored an off-topic guardrail methodology for teams that do not yet have real application data. The method defines acceptable and unacceptable requests, generates varied examples with a language model, and trains classifiers to assess requests against the application’s system instructions. A narrowly scoped service needs to recognize whether a question belongs within its remit, even when the question itself is benign.
Localized moderation with LionGuard 2: Chua co-authored LionGuard 2 with Leanne Tan, Ziyu Ge, and Roy Ka-Wei Lee. The system combines pretrained OpenAI embeddings with a small classifier that assigns severity levels across multiple content categories. It supports English, Chinese, Malay, and partial Tamil and was deployed within Singapore’s government. Multilingual representations and localized training data let the team build moderation without fine-tuning a large language model.
Multilingual evaluation with RabakBench: With Tan, Ge, and Lee, Chua co-developed RabakBench, which tests safety in Singlish, Chinese, Malay, and Tamil. Its construction combines adversarial generation, assisted labeling, translation that preserves harmful meaning, and human validation. The team’s evaluations found weaknesses in existing guardrails across Singapore’s linguistic setting. The contribution includes a method for constructing tests around language that standard evaluations may poorly represent.
Child-specific risks with MinorBench: Chua, Shaun Khoo, and Rachel Shong co-authored MinorBench, informed by a study of an LLM chatbot used in a middle school. They developed categories of content risks for children and tested whether models refused unsafe or inappropriate requests. The intended user becomes part of the evaluation, so a student-facing system can be assessed against risks that an adult-oriented benchmark might miss.
False alarms as well as missed harms: In their automated red-teaming work, Chua and Tan tested both harmful text that escaped moderation and benign text that was incorrectly blocked. An attacking model revised examples using guardrail feedback, while a critic model helped refine its attempts. Chua’s writing on guardrail design also favors configurable scores and responses—logging, warning, modifying, or blocking—so teams can match safeguards to an application’s risks and user experience.
Making AI applications easier to inspect
Chua’s open-source projects give developers concrete ways to explore language-model applications. RAGxplorer visualizes retrieval-augmented generation, in which an application retrieves document content to help answer a question. Developers can load a PDF and visualize queries against its contents, making the relationship between a question and the material available to answer it easier to inspect.
Open NotebookLM adapts knowsuchagency’s existing pdf-to-podcast project into an application that turns a PDF into conversational audio using open-source language and speech models. Its document-to-dialogue-to-MP3 workflow exposes the stages behind the finished audio, giving developers an implementation they can examine and adapt while preserving the underlying project’s collaborative lineage.
These projects complement Chua’s safety research and community work: they give practitioners something tangible to test, understand, and build upon. At OpenAI, his focus is helping developers turn that experimentation into working software with Codex.
Charlie Guo and Gabriel Chua build a live translation app from a Slack conversation, then connect context, browser control, goals, threads, hooks and automations into a repeatable development process.
App shots carry visible application context and embedded text into a task; plugins add reusable practices, service connections and tools.
Choose browser and application control by the task: deep local inspection, operation of software without an API, or access to an existing signed-in browser session.
Long-running goals need verifiable completion criteria, visible progress and a way to surface blockers. A technically satisfied instruction can still miss the intended result.
Subagents separate work for the model; visible threads make workstreams inspectable by the person. Remote threads can delegate browser testing to a local thread and receive findings back.
A repeatable development loop connects incoming feedback to implementation, testing and review, with human input retained where access or judgment requires it.