Portrait of Eungi Hong

Hi, I’m Eungi.

I’m researching coding agent frameworks and benchmarks at NUS, focusing on how to make agents code more strategically for long-horizon and repository-scale tasks. For most of this year, I’ve been working at an early-stage medtech startup, building our outreach and campaign AI for the healthcare and social care space!

Outside of work, I like making fun things happen and bringing people together in my community. If I’m not bringing my company out to ice-skate, then I’m probably planning a blind date night for NUS students with s/unny.

What I’m working on

Three things have my attention right now.

  1. 01

    Coding agent frameworks and benchmarks

    My research at NUS. I’m looking at how frameworks and benchmarks shape whether an agent codes strategically on long-horizon, repository-scale tasks, rather than just reacting to whatever is in front of it.

  2. 02

    Patient engagement AI at Marymount Labs

    I’ve spent most of this year building our outreach and campaign AI for the healthcare and social care space.

  3. 03

    Getting people in a room together

    s/unny, blind date nights for NUS students, and the occasional company ice-skating trip.

Research log

What I’ve been working through.

Notes from my research at NUS. Some entries are analyses of work other people have published, others are my own working notes.

  1. SlopCodeBench

    Orlanski et al., SlopCodeBench: Benchmarking How Coding Agents Degrade Over Long-Horizon Iterative Tasks, March 2026. arXiv:2603.24755

    My summary of the paper, and one of many literature reviews we are doing for our research project on strategic agents.

    Read my notes

    Existing benchmarks don’t test iterative coding

    • They evaluate models against complete, fixed task specifications rather than evolving ones. But in the real world, requirements change over time and code must remain extensible.
    • Under iterative editing, LLMs are known to produce low-quality, high-volume code “slop” that is harder to extend but not penalized by one-shot benchmarks. Two recognizable symptoms:
      • Verbose constructions
      • Redundant methods
    • Recent benchmarks focused on multi-turn or long-horizon coding still fall short:
      • Some construct “iterative” tasks by decomposing a large problem into subproblems; others derive tasks from real open-source commit histories. Both approaches hand the agent a preset specification path, so it is never actually forced to make open-ended architectural decisions.
      • The real test of iterative ability is whether an agent has to build on its own prior code and live with the consequences of its own past architectural choices.

    Introducing SlopCodeBench

    • A benchmark for measuring how code quality evolves as agents repeatedly extend their own prior code under changing specifications. Uses 36 problems and 196 checkpoints.
    • Every problem follows strict design principles to guarantee genuinely open-ended architectural choices:
      • No prescribed internal interfaces. Only external contracts (CLI arguments, API I/O) are specified.
      • No visible test suite. Agents see only specification prose and embedded examples.
      • Black-box, language-agnostic design. Problems constrain only observable behavior.
    • Each problem is an ordered sequence of checkpoints. At every checkpoint, the agent gets a new spec but no memory of the prior conversation. It must make design choices purely based on the current code structure.
    • Code quality is scored via two curated, trajectory-level failure modes, tracked independently of functional correctness. (Existing quality metrics are not used as they are either too generic for iterative development or too narrowly scoped to one specific issue.)
      • Structural erosion. The concentration of control-flow complexity. Suitable because erosion in iterative development is mostly driven by piling haphazard edits onto existing functions rather than writing new ones. Measured as the fraction of the codebase’s total complexity that sits in its high-complexity functions.
      • Verbosity. Verbose, duplicated, and unnecessary code. Measured with 2 metrics:
        • Targeted AST-Grep rules based on observed anti-patterns and best practices
        • Structural code duplication
    • Each checkpoint is graded by running the solution as a subprocess or served API and checking its output against hidden tests.

    Findings from applying SlopCodeBench to frontier models

    • Iteration does not improve correctness (but raises costs). Notably, no agent solves any problem end-to-end.
    • Agent code degrades as it iterates:
      • Structural erosion rises in 77% of trajectories, and verbosity rises in 75.5%.
      • Verbosity’s growth comes mostly from rewriting existing code into duplicates, not from fresh violations.
    • Agent code is more structurally eroded and verbose, and also degrades faster, than human-written code.
      • Agent code is 2.3× more verbose and 2.0× more eroded than open-source repositories. Per checkpoint, agent verbosity grows 6.6× and erosion grows 5.0× faster than in human code.
      • Notably, agents exhibit significantly faster degradation compared to pre-2024 commits.
    • Prompts can reduce initial structural erosion but do not stop iterative degradation.
      • Degradation continues at ~1.3pp/checkpoint even under the best prompt.
      • Only one model, GPT 5.4, showed quality actually improving from start to end.

    Closing thoughts

    • Interesting take, grounded in coding-agent behaviors most of us have run into firsthand. The open question is how to turn this framework into something that actually improves coding agents: what’s a suitable time horizon for iterative feedback, and are the erosion/verbosity metrics used here the right proxies for it?
    • Odd that the benchmark is designed to be language-agnostic and even tests agents on adding support for new languages, given how rare that specific task is in most day-to-day industry work.
  2. A Philosophy of Software Design

    My summary of ‘A Philosophy of Software Design’ by John Ousterhout. My notes are a work in progress and best effort at abstracting key concepts related to ‘complexity’ to eventually concretise our research direction in creating agents that can tackle complexity and code strategically.

    Read my notes

    1. Introduction

    • Software engineering is a creative endeavour in that it is bounded only by our own thought process. This makes our ability as software developers to understand our software system the biggest limitation in software engineering.
    • However, software becomes more complicated as it evolves over time and this slows development. While the accumulation of complexity is inevitable, we must design software more simply to build larger, more powerful systems.
    • Generally, the approaches to tackle complexity are
      • Eliminating it by making code simpler and more obvious.
      • Encapsulating it so that only a small amount is exposed per task. Called modular design.
    • Software design is a continuous process that spans the lifecycle of a software system.
      • Previously, the design was concentrated at the beginning of the project, as per the waterfall model. A software project is divided into discrete phases, and the entire system is designed at once during the design phase. Under this model, when problems arise during implementation, design changes cannot be made easily, and instead patch-around solutions lead to an explosion in complexity.
      • Under agile development, which is more common today, software is designed incrementally. With each iteration, a small subset of functionality is designed and implemented. Problems are discovered, fixed, and improve the next iteration. This incremental approach works because software is malleable; it allows significant design changes partway through implementation.
      • Under this approach, software developers must continuously redesign and improve the design.
    • Since software developers must always think about design issues, and reducing complexity is the most important element in software design, software developers must always think about complexity.

    2. The Nature of Complexity

    Defining complexity
    • The book defines “complexity” as: anything related to the structure of a software system that makes it hard to understand and modify the system.
      • If a software system is hard to understand and modify, it is complicated.
      • If it is easy to understand and modify, it is simple.
    • In terms of cost and benefit,
      • In a complex system, it takes a lot of work to implement even small improvements.
      • In a simple one, larger improvements can be made with less effort.
    • Complexity is what a software developer experiences at a particular point in time when trying to achieve a particular goal, not the overall size and functionality of the system.
    • Complexity is determined by the most common activities.
      • Mathematically, the overall complexity of a system is determined by the complexity of each part weighted by the fraction of time developers spend working on that part.
    • Complexity is more apparent to (and determined by) readers rather than writers.
    • Complexity is incremental, it accumulates in many dependencies and obscurities.
    Symptoms of complexity
    • Change amplification: a simple change requires code modification in many different places.
      • A design goal should be to reduce the amount of code affected by each design decision.
    • Cognitive load: how much a developer must know to complete a task. Higher cognitive load means developers spend more time learning information, and there is greater risk of bugs because they might miss something.
      • An approach requiring more lines of code may be simpler if it reduces cognitive load.
    • Unknown unknowns: it is not obvious which pieces of code must be modified to complete a task, or what information a developer needs to carry out the task successfully.
      • This manifestation of complexity is the worst. There is no certainty on what to do and whether something will work.
      • One of the most important design goals is for a system to be obvious, the opposite of cognitive load and unknown unknowns.
    Causes of complexity
    • Dependency: a given piece of code cannot be understood and modified in isolation.
      • A design goal is to reduce the number of dependencies and make them simple and obvious.
    • Obscurity: important information is not obvious.
      • Can occur when it is not obvious that a dependency exists, or when there are inconsistencies.

    3. Strategic vs Tactical programming

    • Tactical programming focuses on getting things to work, as quickly as possible. Without consideration for the future, complexities accumulate and it becomes harder to work on the codebase and fix the mess.
      • A tactical tornado is a developer who programmes extremely tactically (and leaves waves of destruction behind).
    • Strategic programming focuses on producing a good long-term design rather than prioritising short term speed. Since most code is written by extending an existing code base, a developer’s job is to facilitate future extensions. Produce great design that also happens to work.
      • Invest time to improve the design, both proactively by finding simple designs, and reactively, by fixing mistakes.
    • Make small investments on a continual basis as ideal design trends emerge over time.
    • While fast-paced start up environments create pressure to programme tactically, it is difficult to code once the system has turned messy and one may need to pay high development cost for the remaining life of the product.
      • Hire quality engineers who care about design to lower cost in the long run.

    Programming strategically produces good designs and is cheaper in the long run. To do so, make continuous small investments in good design.

    4. Modules should be deep

    Modular design
    • In modular design, a software system is decomposed into modules that are relatively independent. While dependencies are inevitable, the goal of modular design is to minimize the dependencies between modules.
    • Each module consists of:
      • Interface. What the module does. Consists of everything that a developer working in a different module must know in order to use a given module.
      • Implementation. How it does it. Consists of the code that carries out the promises made by the interface.
    • A module is good when its interface is much simpler than its implementation, because:
      • A simple interface minimises the complexity the module imposes on the rest of the system.
      • There will be many aspects of the module that can be changed without affecting other modules.
    What’s an interface?
    • Consists of:
      • Formal parts. Specified explicitly in code, and can be checked for correctness.
      • Informal parts. High-level behaviour and constraints. If a developer needs to know a particular piece of information to use a module, it is part of the module’s interface.
        • Larger and more complex.
        • Can only be described using comments.
    • A clearly specified interface indicates exactly what developers need to know in order to use that module. This eliminates unknown unknowns.
    What’s an abstraction?
    • “A simplified view of an entity, which omits unimportant details”.
    • In modular programming, each module provides an abstraction in the form of an interface.
    • An abstraction can go wrong in 2 ways:
      • An abstraction includes details that are not actually important. This makes it more complicated and increases cognitive load.
      • An abstraction omits details that are actually important. Developers looking at the abstraction will not have all the information they need to use the abstraction correctly. Hence, this creates obscurity.
        • This is called false abstraction. It looks simple but is not.
    Deep modules
    • Deep modules provide powerful functionality but have simple interfaces. Modules should be deep.
    • Module depth is also about cost and benefit.
      • A module’s benefit is the functionality it provides.
      • A module’s cost is its interface, which represents the complexity it imposes on the rest of the system.
    • Shallow modules are those whose interface is relatively complex compared to the functionality that it provides. The benefit they provide is negated by the cost of learning and using their interfaces. Small modules tend to be shallow.
    Classitis
    • Conventional wisdom that classes should be small, not deep. Minimizes functionality in new classes.
    • Results in individually simple classes, but increases the complexity of the overall system.
    • Interfaces should be designed to make the common case as simple as possible.
      • Instead of dividing functionalities into different classes, make the common case the default.
      • If an interface has many features but most developers only need to be aware of a few of them, the effective complexity of the interface is just the complexity of the commonly used features.

    5. Information hiding

    • Each module should encapsulate a few pieces of knowledge, which represent design decisions. This knowledge is embedded in the module’s implementation but does not appear in its interface, so it is not visible to other modules. Such information includes:
      • Data structures and algorithms.
      • Lower level details.
      • Higher-level abstract concepts (e.g. assumptions).
    • Information hiding reduces complexity in 2 ways:
      • It simplifies the module’s interface, which reduces cognitive load.
      • Makes it easier to evolve the system. Reduces dependencies, since there can’t be dependencies on hidden information. A design change related to the module only affects that module.
    • Hiding more information simplifies the module’s interface and makes the module deeper.
    • Partial information hiding can be valuable too.
      • Limiting access when information is not needed in most common use cases. This creates fewer dependencies than when it is visible to every user.
    Information leakage
    • Occurs when a design decision is reflected in multiple modules. This creates a dependency between modules.
    • Types of information leakage:
      • Leakage through an interface. If a piece of information is reflected in the interface for a module, it is by definition leaked. Hence, simpler interfaces correlate with better information hiding.
      • Back-door leakage. When 2 classes both have knowledge of, and are dependent on, a particular piece of information. Not obvious and thus more pernicious.
    • How to manage?
      • If the affected classes are small and closely tied to the leaked information, merge into a single class.
      • If you can find a simple interface that abstracts away the details, pull the information out of all the affected classes and create a new class that encapsulates just that information.
    Temporal decomposition
    • When the structure of a system corresponds to the time order in which operations will occur.
  3. Formulating complexity

    Where the first two meet. SlopCodeBench scores structural erosion as complexity concentrated in high-complexity functions, while Ousterhout argues for deep modules that put powerful implementations behind simple interfaces. A deep module concentrates complexity on purpose, so the two pull against each other. Working out that tension is the current thrust.

    Read my notes

    The tension

    The book wants deep modules. The paper’s erosion metric penalises exactly the concentration of complexity that a deep module creates. So the open question is whether erosion and verbosity are the right proxies for complexity as Ousterhout defines it, or whether they measure something adjacent to it. For now the book is our standard, and the metrics get compared against it rather than the other way round.

    What we need to figure out

    • Antipatterns. What they actually look like in agent-written code, and how they arise in the first place.
    • Prompts. How they are interpreted, and why prompt-level fixes lower initial erosion without stopping the degradation that follows.
    • Intervention. Where in the loop it is actually possible, and what it would have to change.
    • Evals. What they should measure, if complexity is the thing Ousterhout describes rather than the thing a static score can see.

Want to talk about any of these? Email me.

Projects

Things I’ve built.

Off the clock

Let’s get more personal?

I’m also a performing artist across different dance genres (and dabbling in a cappella!). I find that the thrill of performance is irreplaceable. I also love all things funk: the culture, dance, music & community.

Say hello.

Research, projects, or anything else. I’m easy to reach.