Edward J. SchwartzComputer Security Researcher4 min. read

I previously wrote about OOAnalyzer-ASP, an experimental reimagining of the OOAnalyzer tool using Answer Set Programming. I have bad news: I have stopped working on it for now. I am not sure that anyone besides me will really care, but I'll post this here because I will almost certainly forget the details of why I stopped working on it if I don't write them down.

Performance Problems

The ASP implementation I was using, clingo, is a tight integration of a grounder and a solver. The grounder converts the ASP program into the constraint language expected by the solver. The solver is similar to a SAT solver, but is specialized for Answer Set Programming. One of the reasons I was interested in using ASP is that it uses a concept from SAT, conflict-driven clause learning (CDCL), to prune the search space. This is something that is sorely missing in OOAnalyzer, which only explores the search space to a depth of one. I was hoping that CDCL would be able to efficiently tease out facts that had to be true in order for the program to be consistent.

Grounding

Sadly, there were a few problems with this. The first problem is the grounding process itself. To oversimplify, grounding is a process that removes variables from logic programs. Let's say that we have the following logic program:

foo(ed).
bar(X) :- foo(X).

The grounded version of this program would be:

foo(ed).
bar(ed) :- foo(ed).

The challenge with grounding is that it can lead to a combinatorial explosion in the size of the grounded program. And to make a long story short, because class identities in OOAnalyzer are variable, each potential class identity would amount to grounded facts. This was a constant tension, as I was often thinking about the impact on grounding when I was designing the ASP rules.

Large Lemmas

The second performance problem was related to the solver. I found that it worked very well for small programs, but for medium-sized programs, the solver would often stop making progress. It was possible to eek out better solutions by tuning parameters, but it always seemed like it was only exploring a limited portion of the search space.

As clingo runs, it tries to solve a formula saying "find a model that satisfies all of these constraints, and has a model reward/score of at least X". Once X reaches a certain magnitude, this becomes very difficult to solve because a small portion of the search space will score at least X. More problematically, the conditions required to reach X are very complex and interdependent. This too is very impacted by OOAnalyzer's dynamic class identities. At the solver level, when it finds a portion of the search space that is unsatisfiable, it mechanically generates a lemma that describes it to prevent the solver from exploring that portion of the search space again.

In OOAnalyzer-ASP, these lemmas were often huge, which means that they were very specific to the particular model being solved. At a high level, they might say that "if you have a class consisting of exactly these methods, then you can't have a reward of X", rather than a more general lemma that says "if method A and B are on different classes, you can't have a reward of X". What is supposed to happen is that over time, the lemmas accumulate and new lemmas become smaller and more general, until the lemmas cover the entire search space.

The problem is that solvers remove lemmas over time to speed up the search process. At a certain point, this is at tension with the overall learning process. If lemmas are removed before the smaller, more general lemmas are learned, then the solver will have to re-learn the same lemmas over and over again. It stops making progress. This is what was happening in OOAnalyzer-ASP. I think it is probably possible to tune the solver to avoid this, but the next problem in OOAnalyzer-ASP was more pressing.

Articulation Challenges

The second major problem with OOAnalyzer-ASP was that it was difficult to adequately express some properties in ASP. More specifically, as I was fighting the above performance problems, I would often convince the solver to find a model with a better score. Unfortunately, when I actually compared the better scoring model with the ground truth, it would often be slightly worse. It was very difficult to actually construct the model scoring to reflect reality.

One of the reasons for this has to do with fixed points. OOAnalyzer represents classes as sets of methods that are iteratively merged. In Prolog, it is possible to reason about the membership of the current class at each step of the iteration. For example, OOAnalyzer has a few rules that are conditioned on one class only containing a single class. In ASP, class membership is represented as a fixed point, which means that class properties must be monotonic over merging. There is not an easy way to express the fact that a class only contains a single method at some point in time, because the rules can only reason about the final fixed point.

Imagine you had a program like this:

merge(A,B) :- class(A), class(B), hasAConstructor(A), hasADestructor(B), size(B) = 1.

In ASP, after merging A and B, size(B) = 1 would become false, and this line of reasoning would vanish.

In theory, it is possible to represent the entire history of the merging process in ASP by modeling the entire sequence of merges, but this would further compound the grounding performance problem!

Conclusion

I don't think that OOAnalyzer-ASP is a complete dead-end, but it's not the slam dunk I was hoping it would be.

As part of this process, I also spent a lot of time looking at the mistakes OOAnalyzer and OOAnalyzer-ASP made. There is less room for improvement than I expected. That doesn't mean that OOAnalyzer is perfect, but rather, after looking at the available evidence, I am not sure how as a human to do much better. I honestly expected that OOAnalyzer's limited searching ability would have a larger impact than it does.

In Part 1 and Part 2, I reported some peculiar behavior for quantized Qwen3.5-2B models. In Part 3, I reported results for a larger variant, Qwen3.5-35B-A3B, with thinking enabled.

In this post, we'll take a look at performance with thinking enabled vs. disabled.

Sweep Summary

RunResolved%PPLKLRuntimeExceptions
thinking-BF16173/50034.6%6.621392m 16sTimeout(15), ExitCode(22), Reward(2), Verifier(1)
thinking-Q5_K_M180/50036.0%6.620.00831046m 35sTimeout(8), ExitCode(19), Reward(5)
thinking-Q8_0193/50038.6%6.610.00681646m 39sSetup(7), Timeout(12), ExitCode(18), Reward(4), Verifier(1)
thinking-vllm233/50046.6%1753m 3sTimeout(41), NetworkConnectionError(1), ExitCode(28), Reward(3), Verifier(1)
nonthinking-BF16273/50054.6%6.62-0.00001104m 26sTimeout(1), ExitCode(22), Reward(2), Verifier(1)
nonthinking-Q8_0260/50052.0%6.610.0068890m 2sTimeout(2), ExitCode(18), Reward(2), Verifier(1)
nonthinking-Q5_K_M263/50052.6%6.620.0083877m 18sTimeout(2), ExitCode(23), Reward(6), Verifier(1)
nonthinking-vllm269/50053.8%734m 0sTimeout(4), ExitCode(33), Reward(1), Verifier(1)

Analysis

As expected, the thinking runs took longer to complete than the non-thinking runs. But unexpectedly, non-thinking runs outperformed thinking runs across the board. This is counterintuitive, since we would expect that thinking would improve performance. That is how it's supposed to work!

One concern I had with this experiment was whether timeouts were affecting the results, that is, if given enough time, the thinking runs would eventually outperform the non-thinking runs. I actually ended up running this several times, eventually with quite long timeouts. Most of the remaining timeouts were due to thinking loops rather than legitimate reasoning timeouts. So endless loops are a concern with this particular model on thinking runs. But we see more differences in performance --- exactly 100 more problems were resolved by nonthinking-BF16 than thinking-BF16 --- that can't be explained by the 15 observed timeouts in the thinking-BF16 run.

I asked AI to analyze the results and it found two problems.

Problem 1: Tool-call parsing

In thinking mode, llama.cpp is experiencing a malformed tool-call parse crash that is causing a large number of failures. These failures cause openhands to exit.

I can trigger this behavior using this script:

=== firing the same clean real prompt 10 times ===
POST http://172.17.0.2:8085/v1/chat/completions x 10

  probe 1: EXTRACTED (str_replace_editor)
  probe 2: EXTRACTED (str_replace_editor)
  probe 3: EMPTY (finish_reason=stop, no content, no reasoning)
  probe 4: EMPTY (finish_reason=stop, no content, no reasoning)
  probe 5: CRASH (500: Failed to parse input at pos 49: <tool_call>
<function=str_replace_editor>
<parameter=command>
view
)
  probe 6: EMPTY (finish_reason=stop, no content, no reasoning)
  probe 7: EMPTY (finish_reason=stop, no content, no reasoning)
  probe 8: CRASH (500: Failed to parse input at pos 49: <tool_call>
<function=str_replace_editor>
<parameter=command>
view
)
  probe 9: EMPTY (finish_reason=stop, no content, no reasoning)
  probe 10: EMPTY (finish_reason=stop, no content, no reasoning)

  -> extracted: 2/10, stuck: 0/10, crash: 2/10, empty: 6/10, other: 0/10

So out of 10 queries, we get the failure to parse twice, and then an empty response 6 times. Something is clearly going wrong!

With the help of AI, I was able to attribute many of the problems to a grammar problem. The model often emits a redundant <thinking> tag, which causes llama.cpp's parser to fail. As you can see in this AI-generated table, this happens very frequently:

ModelResolvedHarbor exception¹Parse-crash²Stuck-loop³Other silent-crash⁴Unexplained⁵
BF161733514485162
Q8_01933999107161
Q5_K_M1803196135355
vLLM233650911398

¹ Timeout/ExitCode/Setup/Reward/Verifier/NetworkConnectionError — the only category Harbor's own exit-code-based detection actually catches. ² llama.cpp's PEG parser crash (Failed to parse input at pos N), trial ended fatally right there. Zero on vLLM. ³ OpenHands' AgentStuckInLoopError killed the attempt. ⁴ Other unrelated fatal errors (mostly the chardet/action-execution-server bug on vLLM). ⁵ Attempts that neither resolved, hit a Harbor exception, nor showed any of the fatal patterns checked — genuinely completed but produced a wrong patch, or a crash pattern not yet characterized.

Problem 2: Harbor error handling

The second problem is that Harbor does not detect the error. And it seems that this is because openhands returns exit code 0.

What a mess!

Concluding Thoughts

We can't detect the redundant <thinking> tag problem in vllm, but it could still be happening (update: it is happening). vllm has a different tool parsing mechanism than llama.cpp; it might be more forgiving (update: it is). For whatever reason, vllm-thinking is still underperforming vllm-nonthinking.

According to Qwen's blog post, Qwen3.5-35B-A3B scores 70.0% on SWE-bench Verified. There is a small footnote too:

SWE-Bench Series: Internal agent scaffold (bash + file-edit tools); temp=1.0, top_p=0.95, 200K context window. We correct some problematic tasks in the public set of SWE-bench Pro and evaluate all baselines on the refined benchmark.

I'm not going to comment on the "correct some problematic tasks" part.

That temperature of 1.0 and top_p of 0.95 are recommended by Qwen for both thinking "general tasks" (as opposed to "precise coding tasks") and non-thinking "reasoning tasks" (as opposed to "general tasks"). Unfortunately, as you can see here, for these results I used "precise coding tasks" and "general tasks", which both use different temperature and top_p settings. If you are wondering why SWE-Bench Verified is not considered a "precise coding task"... well, so am I. Ask Qwen! IMHO, models should have a single set of recommended parameters for all tasks, and the model should be able to handle the task type automatically.

Qwen's claimed performance on SWE-bench Verified
Qwen's claimed performance on SWE-bench Verified

Next Steps

  1. I am going to try to confirm whether the redundant <thinking> tag is generated by vllm-thinking. If it is, then it suggests that the model is buggy. Which would be weird --- the model defaults to thinking. (Update: I confirmed that vllm-thinking does generate the redundant <thinking> tag, so its parser must be more forgiving than llama.cpp's parser.)

  2. I am going to rerun the experiments with all four parameter settings that Qwen recommends, since I apparently chose different ones than Qwen did when they tested SWE-Bench Verified.

  3. I'm will test the dense model Qwen3.5-27B.

  4. I am also going to investigate agents other than openhands. I originally selected openhands because it is the first agent I was able to get to work with Harbor. But at the time, Harbor was only compatible with an older branch of openhands. Hopefully this has changed, or a more reasonable harness like opencode will work now.

In Part 1 and Part 2, I reported some peculiar behavior for quantized Qwen3.5-2B models. In Part 3 Take 1, I reported early results for a larger variant: Qwen3.5-35B-A3B. But I recently discovered that this experiment had a problem that may have been contributing to Timeouts.

The Issue

Qwen3.5 can be used in either "thinking" or "non-thinking" mode. Small models like Qwen3.5-2B default to non-thinking mode, but larger models like Qwen3.5-35B default to thinking mode. When I switched to testing Qwen3.5-35B-A3B, I did not realize that the default mode had changed. That's the first problem; we were testing Qwen3.5-2B without thinking against Qwen3.5-35B-A3B with thinking. The second problem is that Qwen recommends different sampling parameters for thinking vs. non-thinking mode. So we were using a non-thinking sampling configuration with a thinking model. This is likely to have contributed to the Timeouts we observed in Take 1. Manual analysis revealed that many of these cases were indeed thinking loops, which is consistent with the fact that we were using an invalid sampling configuration with a thinking model.

So, mea culpa. Fortunately, I noticed this problem because I wondered why vllm had so many Timeouts and investigated further. I think this is a good example of why we need to conduct and automate these experiments. It's easy to make mistakes when running these experiments manually, and automation can help catch these issues, or at least make sure we don't repeat them once we noticed the problem. And even if you aren't running experiments, it's easy to accidentally use a model with unsupported parameters and get poor results. So it's important to understand the model and make sure you are running with all of the correct parameters.

Take Two

So let's try this again but explicitly set the model to thinking mode and use the recommended sampling parameters for thinking. Recall that we were attempting to see whether the trends we observed for Qwen3.5-2B would hold for Qwen3.5-35B-A3B:

  1. KL divergence did not reliably predict downstream coding-agent performance.
  2. Some quantizations outperformed the original model, while others underperformed.

Old, Invalid Results

Here was the old, invalid run summary from Take 1:

RunResolved%PPLKLRuntimeExceptions
BF16181/50036.2%6.622103m 41sTimeout(275), ExitCode(17), Reward(3), Verifier(1)
Q5_K_M194/50038.8%6.620.0083531m 46sTimeout(8), ExitCode(23), Reward(3), Verifier(1)
Q8_0202/50040.4%6.610.0068575m 45sTimeout(11), ExitCode(24), Reward(1), Verifier(1)
vllm230/50046.0%1634m 13sTimeout(145), ExitCode(22), Reward(6)

And here were the key takeaways from Take 1:

  1. vllm currently leads this sweep at 46.0% resolved.
  2. Q8_0 and Q5_K_M both outperform the GGUF BF16 variant in resolved rate.
  3. BF16 and vllm show many more timeouts than the quantized GGUF variants.
  4. Public benchmark numbers for this model are meaningfully higher than what I observe locally.

New, Corrected Results

Let's look at the corrected (hopefully) results from Take 2. Here is the new run summary:

RunResolved%PPLKLRuntimeExceptions
BF16174/50034.8%6.621451m 33sAddTestsDirError(1), Timeout(24), ExitCode(20), Reward(2)
Q5_K_M187/50037.4%6.620.0083978m 48sTimeout(11), ExitCode(19), Reward(1), Verifier(1)
Q8_0181/50036.2%6.610.0068989m 28sTimeout(7), ExitCode(24), Reward(1)
vllm249/50049.8%1605m 47sTimeout(42), ExitCode(19), Reward(5), Verifier(2)

New takeaways:

  1. vllm still leads this sweep at 49.8% resolved.
  2. Q5_K_M and Q8_0 still outperform the GGUF BF16 variant in resolved rate.
  3. BF16 and vllm still show more timeouts than the quantized GGUF variants, but the gap is much smaller than in Take 1. Manual analysis showed that the models were still making progress rather than being stuck in loops.
  4. Public benchmark numbers for this model are meaningfully higher than what I observe locally.

If you look closely, the takeaways are essentially the same. Oddly enough, vllm improved its performance with the correct configuration, while the GGUF variants all performed worse.

Observations

Behavior differs from Qwen3.5-2B

With Qwen3.5-2B, the smaller GGUF quantizations (Q5_K_M, Q8_0) outperformed both GGUF BF16 and the vllm model. Qwen3.5-35B-A3B shows a similar pattern, but the vllm model outperforms all the GGUF variants.

GGUF BF16 vs original checkpoint is still surprising

There is a notable gap between GGUF BF16 results and the original model results. That is surprising because the original checkpoint is mostly BF16 with a small number of F32 tensors. I checked the GGUF contents and confirmed F32 tensors are present:

vscode ➜ /workspaces/auto-bench (main) $ gguf-dump /home/vscode/.cache/huggingface/hub/models--unsloth--Qwen3.5-35B-A3B-GGUF/snapshots/bc014a17be43adabd7066b7a86075ff935c6a4e2/BF16/Qwen3.5-35B-A3B-BF16-00002-of-00002.gguf | grep "F32" | head -20
INFO:gguf-dump:* Loading: /home/vscode/.cache/huggingface/hub/models--unsloth--Qwen3.5-35B-A3B-GGUF/snapshots/bc014a17be43adabd7066b7a86075ff935c6a4e2/BF16/Qwen3.5-35B-A3B-BF16-00002-of-00002.gguf
    142:       2048 |  2048,     1,     1,     1 | F32     | blk.0.attn_norm.weight
    143:         32 |    32,     1,     1,     1 | F32     | blk.0.ssm_a
    144:      32768 |     4,  8192,     1,     1 | F32     | blk.0.ssm_conv1d.weight
    145:         32 |    32,     1,     1,     1 | F32     | blk.0.ssm_dt.bias
    148:        128 |   128,     1,     1,     1 | F32     | blk.0.ssm_norm.weight
    149:     524288 |  2048,   256,     1,     1 | F32     | blk.0.ffn_gate_inp.weight
    153:       2048 |  2048,     1,     1,     1 | F32     | blk.0.ffn_gate_inp_shexp.weight
    154:       2048 |  2048,     1,     1,     1 | F32     | blk.0.post_attention_norm.weight
    155:       2048 |  2048,     1,     1,     1 | F32     | blk.1.attn_norm.weight
    156:         32 |    32,     1,     1,     1 | F32     | blk.1.ssm_a
    157:      32768 |     4,  8192,     1,     1 | F32     | blk.1.ssm_conv1d.weight
    158:         32 |    32,     1,     1,     1 | F32     | blk.1.ssm_dt.bias
    161:        128 |   128,     1,     1,     1 | F32     | blk.1.ssm_norm.weight
    162:     524288 |  2048,   256,     1,     1 | F32     | blk.1.ffn_gate_inp.weight
    166:       2048 |  2048,     1,     1,     1 | F32     | blk.1.ffn_gate_inp_shexp.weight
    167:       2048 |  2048,     1,     1,     1 | F32     | blk.1.post_attention_norm.weight
    168:       2048 |  2048,     1,     1,     1 | F32     | blk.10.attn_norm.weight
    169:         32 |    32,     1,     1,     1 | F32     | blk.10.ssm_a
    170:      32768 |     4,  8192,     1,     1 | F32     | blk.10.ssm_conv1d.weight
    171:         32 |    32,     1,     1,     1 | F32     | blk.10.ssm_dt.bias

In fact, some tensors appear upcast to F32 in the GGUF file, so raw dtype alone does not explain the performance gap.

Vendor-reported benchmark numbers are much higher

According to Qwen's benchmark page, Qwen3.5-35B-A3B achieves 69.2% on SWE-bench Verified. That is substantially higher than the 49.8% best result in this sweep.

Interestingly, on the Qwen3-Coder-Flash model page, Qwen reports 51.6% on SWE-bench Verified using OpenHands, the same agent framework I am using.

Qwen3-Coder-Flash benchmarks
Qwen3-Coder-Flash benchmarks

There is also active community discussion about reproducibility for these numbers here. I have not yet found detailed methodology documentation for Qwen3.5's SWE-bench Verified evaluation setup, which may explain part of the discrepancy.

Next Steps / Questions

  1. Run Qwen3.5-35B-A3B in non-thinking mode.
  2. Test other Qwen3.5 models.
  3. Test non-Qwen models.
  4. Why does BF16 perform worse than the original checkpoint?
  5. Switch from time-based timeouts (the default in Harbor) to iteration-based limits.

This post contains a broken experiment in which thinking was inadvertently enabled but did not use sampling parameters intended for thinking. See here for a corrected version.

In Part 1 and Part 2, I reported some peculiar behavior for quantized Qwen3.5-2B models:

  1. KL divergence did not reliably predict downstream coding-agent performance.
  2. Some quantizations outperformed the original model, while others underperformed.

This post shares early results for a larger variant: Qwen3.5-35B-A3B.

Key Takeaways

  1. vllm currently leads this sweep at 46.0% resolved.
  2. Q8_0 and Q5_K_M both outperform the GGUF BF16 variant in resolved rate.
  3. BF16 and vllm show many more timeouts than the quantized GGUF variants.
  4. Public benchmark numbers for this model are meaningfully higher than what I observe locally.

Sweep Summary

RunResolved%PPLKLRuntimeExceptions
BF16181/50036.2%6.622103m 41sTimeout(275), ExitCode(17), Reward(3), Verifier(1)
Q5_K_M194/50038.8%6.620.0083531m 46sTimeout(8), ExitCode(23), Reward(3), Verifier(1)
Q8_0202/50040.4%6.610.0068575m 45sTimeout(11), ExitCode(24), Reward(1), Verifier(1)
vllm230/50046.0%1634m 13sTimeout(145), ExitCode(22), Reward(6)

Observations

Timeout behavior is a major differentiator

BF16 and vllm both have substantial timeout counts. At this point, it is unclear whether these are true long-horizon failures or loop-like behaviors similar to what I saw with weaker Qwen3.5-2B quantizations. Either way, timeout behavior appears to be a key driver of aggregate score differences.

Behavior differs from Qwen3.5-2B

With Qwen3.5-2B, the smaller GGUF quantizations (Q5_K_M, Q8_0) outperformed both GGUF BF16 and the vllm model. Qwen3.5-35B-A3B shows a similar pattern, but the vllm model outperforms all the GGUF variants.

GGUF BF16 vs original checkpoint is still surprising

There is a notable gap between GGUF BF16 results and the original model results. That is surprising because the original checkpoint is mostly BF16 with a small number of F32 tensors. I checked the GGUF contents and confirmed F32 tensors are present:

vscode ➜ /workspaces/auto-bench (main) $ gguf-dump /home/vscode/.cache/huggingface/hub/models--unsloth--Qwen3.5-35B-A3B-GGUF/snapshots/bc014a17be43adabd7066b7a86075ff935c6a4e2/BF16/Qwen3.5-35B-A3B-BF16-00002-of-00002.gguf | grep "F32" | head -20
INFO:gguf-dump:* Loading: /home/vscode/.cache/huggingface/hub/models--unsloth--Qwen3.5-35B-A3B-GGUF/snapshots/bc014a17be43adabd7066b7a86075ff935c6a4e2/BF16/Qwen3.5-35B-A3B-BF16-00002-of-00002.gguf
    142:       2048 |  2048,     1,     1,     1 | F32     | blk.0.attn_norm.weight
    143:         32 |    32,     1,     1,     1 | F32     | blk.0.ssm_a
    144:      32768 |     4,  8192,     1,     1 | F32     | blk.0.ssm_conv1d.weight
    145:         32 |    32,     1,     1,     1 | F32     | blk.0.ssm_dt.bias
    148:        128 |   128,     1,     1,     1 | F32     | blk.0.ssm_norm.weight
    149:     524288 |  2048,   256,     1,     1 | F32     | blk.0.ffn_gate_inp.weight
    153:       2048 |  2048,     1,     1,     1 | F32     | blk.0.ffn_gate_inp_shexp.weight
    154:       2048 |  2048,     1,     1,     1 | F32     | blk.0.post_attention_norm.weight
    155:       2048 |  2048,     1,     1,     1 | F32     | blk.1.attn_norm.weight
    156:         32 |    32,     1,     1,     1 | F32     | blk.1.ssm_a
    157:      32768 |     4,  8192,     1,     1 | F32     | blk.1.ssm_conv1d.weight
    158:         32 |    32,     1,     1,     1 | F32     | blk.1.ssm_dt.bias
    161:        128 |   128,     1,     1,     1 | F32     | blk.1.ssm_norm.weight
    162:     524288 |  2048,   256,     1,     1 | F32     | blk.1.ffn_gate_inp.weight
    166:       2048 |  2048,     1,     1,     1 | F32     | blk.1.ffn_gate_inp_shexp.weight
    167:       2048 |  2048,     1,     1,     1 | F32     | blk.1.post_attention_norm.weight
    168:       2048 |  2048,     1,     1,     1 | F32     | blk.10.attn_norm.weight
    169:         32 |    32,     1,     1,     1 | F32     | blk.10.ssm_a
    170:      32768 |     4,  8192,     1,     1 | F32     | blk.10.ssm_conv1d.weight
    171:         32 |    32,     1,     1,     1 | F32     | blk.10.ssm_dt.bias

In fact, some tensors appear upcast to F32 in the GGUF file, so raw dtype alone does not explain the performance gap.

Vendor-reported benchmark numbers are much higher

According to Qwen's benchmark page, Qwen3.5-35B-A3B achieves 69.2% on SWE-bench Verified. That is substantially higher than the 46.0% best result in this sweep.

Interestingly, on the Qwen3-Coder-Flash model page, Qwen reports 51.6% on SWE-bench Verified using OpenHands, the same agent framework I am using.

Qwen3-Coder-Flash benchmarks
Qwen3-Coder-Flash benchmarks

There is also active community discussion about reproducibility for these numbers here. I have not yet found detailed methodology documentation for Qwen3.5's SWE-bench Verified evaluation setup, which may explain part of the discrepancy.

Next Steps

I think the critical question right now is whether the timeouts are legitimate. After reviewing the logs, I think that they may not be. So I'm going to re-run with fewer agents in parallel.

In Part 1, I introduced auto-bench, a tool for benchmarking quantized LLMs for local coding agents, and shared some results from a preliminary study on a single instance from SWE-bench Verified. The results showed that (1) KL Divergence doesn't predict performance, and (2) quantizations can both outperform and underperform the original model.

In this post, I'll share some new results. Like the other experiment, this one also focuses on Qwen3.5-2B. Unlike the other experiment, which tested a single instance of SWE-bench Verified with eight attempts, this experiment tests all instances of SWE-bench Verified with one attempt.

Results

Without further ado, here are the results.

QuantResolved%PPLKLRuntimeExceptions
BF1628/5005.6%13.38547m 59sTimeout(47), ExitCode(22), Verifier(2)
IQ4_NL26/5005.2%13.670.0309423m 21sTimeout(29), ExitCode(16), Verifier(1)
IQ4_XS24/5004.8%13.680.0318493m 7sTimeout(39), ExitCode(15), Reward(1), Verifier(2)
Q3_K_M30/5006.0%14.330.0774785m 43sTimeout(67), ExitCode(21)
Q3_K_S20/5004.0%15.080.1334742m 25sTimeout(73), ExitCode(25), Reward(1), Verifier(1)
Q4_024/5004.8%13.910.0454407m 34sTimeout(25), ExitCode(21), Verifier(1)
Q4_136/5007.2%13.680.0273766m 19sTimeout(59), ExitCode(16), Reward(1), Verifier(1)
Q4_K_M27/5005.4%13.790.0230357m 0sTimeout(20), ExitCode(23), Verifier(2)
Q4_K_S19/5003.8%13.780.0274519m 39sTimeout(38), ExitCode(23), Reward(1), Verifier(1)
Q5_K_M62/50012.4%13.460.0082784m 58sTimeout(61), ExitCode(23), Reward(1), Verifier(1)
Q5_K_S46/5009.2%13.490.0100563m 27sTimeout(30), ExitCode(25), Reward(1), Verifier(3)
Q6_K58/50011.6%13.480.0035820m 37sTimeout(62), ExitCode(20), Verifier(1)
Q8_037/5007.4%13.390.0012598m 13sTimeout(46), ExitCode(17), Verifier(1)
UD-IQ2_M1/5000.2%17.610.26771866m 13sTimeout(300), ExitCode(24), Verifier(1)
UD-IQ2_XXS1/5000.2%27.110.70182196m 49sTimeout(371), ExitCode(19)
UD-IQ3_XXS5/5001.0%15.310.15491481m 24sTimeout(230), ExitCode(24), Verifier(1)
UD-Q2_K_XL2/5000.4%17.150.2388442m 51sTimeout(29), ExitCode(27)
UD-Q3_K_XL48/5009.6%13.940.0520738m 44sTimeout(57), ExitCode(19), Reward(1), Verifier(1)
UD-Q4_K_XL57/50011.4%13.600.0164759m 13sTimeout(48), ExitCode(18), Reward(2), Verifier(2)
UD-Q5_K_XL62/50012.4%13.510.0077932m 37sTimeout(71), ExitCode(18), Reward(2), Verifier(1)
UD-Q6_K_XL29/5005.8%13.480.0020539m 38sTimeout(35), ExitCode(20), Verifier(1)
UD-Q8_K_XL36/5007.2%13.370.0011502m 1sSetup(1), Timeout(35), ExitCode(23), Verifier(2)
vllm26/5005.2%521m 14sTimeout(49), ExitCode(11), Verifier(1)

And the plot of % Resolved vs KL Divergence:

Percentage instances resolved vs. KL Divergence
Percentage instances resolved vs. KL Divergence

Observations

A few observations:

  1. The overall resolve rates are low across the board. This is not a very powerful model. I intentionally selected an easy problem instance for the Part 1 experiment.
  2. As in Part 1, many of the "mid-range" quantizations outperform the original model, yet small and large quantizations underperform. This is consistent with the idea that some quantizations are actually beneficial, while others are harmful.
  3. Also as in Part 1, KL Divergence does not fully explain performance.

What next?

This experiment largely confirmed the findings from Part 1 about the Qwen3.5-2B model. An open question is whether these results apply to other models as well. I plan to run similar experiments on larger variants of the Qwen3.5 family next, but I won't be evaluating every quantization. Too much time is wasted on bad quantizations because they get stuck in endless loops. Instead, I'll probably try a select few quantizations, such as BF16, Q8_0, and Q5_K_M. Although I am interested in understanding these peculiar behaviors, my primary goal is actually to find which models and quantizations are usable.

Edward J. SchwartzComputer Security Researcher6 min. read

OOAnalyzer is one of my favorite projects for a few reasons. The underlying problem is deceptively hard. If you look closely enough, OO executables have a lot of evidence in them. When you consider each piece in isolation, it can seem simple. But when you try to combine all of the evidence, you wind up with a combinatorial explosion of possibilities that is pruned by constraints in very complex ways. I also like it because it's a very practical problem. People actually use OOAnalyzer, despite all of its limitations—a testament to how even flawed solutions to hard problems can be valuable.

OOAnalyzer is a successor to an earlier project that predated me at SEI called ObjDigger. OOAnalyzer's big innovation was to use what we called "hypothetical reasoning" about ambiguous scenarios. In a nutshell, we often faced a choice between several possibilities. For example, class D inherits from class B, and we see class M in class D's vftable. We know that M is either a method of D or a method of B. We can't tell which one it is, but we can make a guess, and see if that leads to a contradiction. If it does, then we know our guess was wrong, and we can eliminate that possibility and try the others. This is a powerful technique that was partially enabled by our use of Prolog in OOAnalyzer, specifically Prolog's backtracking search.

That being said, OOAnalyzer's hypothetical reasoning is not perfect. One problem we encountered while developing OOAnalyzer was that there could be a long gap between a guess and when the contradiction is detected. Let's say we make a guess that eventually will cause a contradiction, but we must make 10 other boolean guesses before detecting the contradiction. At that point, Prolog would backtrack the guesses, starting with the most recent guess. Unfortunately, Prolog has no way of knowing that the actual cause was much further back, and it would waste time re-exploring 2102^{10} models that were doomed to fail. Over time, we began to shape the rules in OOAnalyzer to avoid this problem. We ordered the rules so that any guess likely to lead to a contradiction would be detected immediately. This worked, but also limited our reasoning power.

Another problem with OOAnalyzer is how it decides on a final model. The final model is simply the first one that allows all guesses (uncertainties) to be resolved and does not lead to a contradiction. OOAnalyzer's guessing rules are ordered so that the more important ones are made earlier, when they have less chance of conflicting with a previous decision. But there is no guarantee that this model is the best one, or even close to it! Over time, I formed the opinion that OOAnalyzer is really an optimization problem. Our rules permit multiple models that could explain the evidence. But some are better than others. For example, many OO programs without vftables and vbtables can be explained by a model with no OO classes or methods. This is valid, but not very useful for the analyst.

Finally, OOAnalyzer involves a lot of constraints. For example, we might know that a method is either a constructor, real destructor, or deleting destructor. Obviously if we learn that the same method is not a destructor, it must be a constructor. Because we didn't use constraints in OOAnalyzer, we had to implement this type of logic manually. It works, but it's not elegant.

Over time, I've wondered how to better frame the OOAnalyzer problem, and recently I started exploring Answer Set Programming (ASP) as a potential answer. ASP is a form of declarative programming that is based on the stable model semantics of logic programming. It is designed to solve combinatorial search problems, and it has built-in support for constraints and optimization. ASP has several key features:

  • It allows for declarative rules, like Prolog.
  • It has built-in support for constraints. 1 { constructor(X); real_destructor(X); deleting_destructor(X) } 1 means that exactly one of the three predicates must be true for any given X.
  • It has built-in support for optimization. We can assign weights to different models and ask the solver to find the global optimum, or the best solution found within a time limit.

In my spare time, I've been porting parts of OOAnalyzer to ASP here. I've been pleasantly surprised by how well it has worked so far. The code is much more concise and easier to read than the Prolog version. The constraints are much easier to express. For example, here's a rule that says if we see certain behavior (like installing vftables), we know that the method is a constructor or destructor:

% A method that writes a vftable/vbtable into its own this-pointer must be
% exactly one of: constructor, real destructor, or deleting destructor.
% Covers:
%   reasonConstructor (rules.pl:192) — elimination: only remaining candidate
%   reasonRealDestructor (rules.pl:388) — elimination: only remaining candidate
%   reasonDeletingDestructor (rules.pl:575) — elimination: only remaining candidate
%!trace_rule {"% is exactly one of constructor, real destructor, or deleting destructor", Method}
1 { constructor(Method) ; realDestructor(Method) ; deletingDestructor(Method) } 1 :-
    certainConstructorOrDestructor(Method).

Similarly, here's a rule saying that a method can only be one of a constructor, real destructor, or deleting destructor:

constructorDestructorKind(Method, constructor) :-
    constructor(Method).
constructorDestructorKind(Method, realDestructor) :-
    realDestructor(Method).
constructorDestructorKind(Method, deletingDestructor) :-
    deletingDestructor(Method).

% Covers:
%   insanityConstructorAndRealDestructor (insanity.pl:226)
%   insanityConstructorAndDeletingDestructor (insanity.pl:240)
%   reasonNOTConstructor_B (rules.pl:281) — realDestructor → notConstructor
%   reasonNOTConstructor_C (rules.pl:288) — deletingDestructor → notConstructor
%   reasonNOTRealDestructor_B (rules.pl:443) — constructor → notRealDestructor
%   reasonNOTRealDestructor_C (rules.pl:448) — deletingDestructor → notRealDestructor
%   reasonNOTDeletingDestructor_B (rules.pl:632) — constructor → notDeletingDestructor
%   reasonNOTDeletingDestructor_C (rules.pl:637) — realDestructor → notDeletingDestructor
%!trace_rule {"% has multiple constructor/destructor kinds", Method}
insanity(insanityMultipleConstructorDestructorKinds, (Method,Count)) :-
    method(Method),
    Count = #count { Kind : constructorDestructorKind(Method, Kind) },
    Count > 1.

As the comment above says, this rule covers about 8 different rules in OOAnalyzer. In ASP, we can express the same thing much more concisely.

OOAnalyzer-ASP can already read OOAnalyzer's facts format. You can see some examples of the results in the repository. Bear in mind that not all rules are ported yet.

One last thing I'll share is automatic explanations of models. I've been using the xclingo2 library to generate explanations for the ASP models. Even as OOAnalyzer's developer, it can be hard to understand how it reasons, because it could involve dozens of related conclusions. This is exacerbated in ASP because constraints and the optimization criteria are "silent". But even so, xclingo2 can generate detailed proof trees showing why a particular atom was included in the model. Here's an example of a proof tree for why a method is a constructor:

  |__4266400 is a constructor
  |  |__4266400 is a method;4266400 is a method because 4266176 calls it at this-offset 0
  |  |  |__thunk 4264566 resolves to 4266400
  |  |  |  |__thunk 4266400 resolves to 4266400
  |  |  |__4266176 is a method;4266176 is a method because it is a known constructor
  |  |  |  |__4266176 is a constructor;4266176 is exactly one of constructor, real destructor, or deleting destructor
  |  |  |  |  |__4266176 must be a constructor or destructor because it writes a vftable into its own this-pointer
  |  |  |  |  |  |__4266176 writes confirmed vftable 4290632 at offset 0
  |  |  |  |  |  |  |__4290632 is a confirmed vftable;4290632 is a confirmed vftable because RTTI says so

So is this the next generation OOAnalyzer? Theoretically, I think the answer is yes! But practically, it's unclear how well this is going to scale to real programs. One of the downsides to most ASP implementations is that they ground eagerly. Because executables can have a lot of evidence, it's possible this can lead to a combinatorial explosion in the number of ground rules. Because OOAnalyzer's rules involve negation and recursion in complex ways, it's hard to say how much of an issue this will be. We're able to find the optimal model on all of the toy OOAnalyzer programs, but that doesn't mean much. There are alternative approaches to ASP, like lazy grounding, but they are less mature. I'm hopeful that with cautious engineering, we can make this work on real programs, but only time will tell. In the meantime, I'm having fun exploring this new approach to the problem!

Edward J. SchwartzComputer Security Researcher2 min. read

Like many people, when I'm on the go and need to do something, I reach for an LLM assistant. In my case, that's Claude. Anthropic, OpenAI, and Google all have a variety of connections to various services. But what happens when they don't have a connection to a service that you want?

For example, I've been dieting, which means I need to track my calories. I use MyFitnessPal for that. Tracking calories is laborious and difficult. Much of the difficulty comes from estimating the calories in a meal. What are the ingredients and components, and how much of each is there? If only there was a way to take a picture and have an intelligent system analyze the image and determine the calories. Oh wait, we do—LLM assistants!

Sadly, as far as I know, there's no connection to MyFitnessPal in any of the major LLM assistants. Wouldn't it be nice if my LLM could automatically log the estimated calories for me after analyzing the image?

This is where self-hosted MCP servers come in. Claude now allows you to add "custom integrations" to your account that it can use on the mobile app or on claude.ai. These custom integrations are just SSE-based MCP servers. If you have a server, you can run these integrations yourself.

There were a few complications for me:

  1. There are several MyFitnessPal MCP servers, but they are stdio servers, which do not expose the HTTP endpoints that Claude's custom integrations require.

  2. Although I have a server, it is behind a firewall, and I don't want to change that.

Fortunately, these problems aren't that difficult to solve. For the first problem, I used supergateway to expose the stdio MCP server as an SSE MCP server. For the second problem, I used ngrok to create a secure tunnel to my server without changing my firewall settings.

Setting up projects like this has never been easier. I asked Claude Code to help, and it generated the docker-compose solution in about thirty minutes.

As is becoming more frequent, I found that the existing MyFitnessPal MCP servers were all lacking features I needed, such as the ability to see recent and favorite foods, delete log entries, and so on. So I forked them and vibe-coded the features I wanted. Forking the MCP servers and customizing them took a couple of hours spread over a few days of testing. Coding agents make this type of project easier than ever.

I'm still working out some edge cases with meal recognition and calorie estimation accuracy, but I've been really happy with the core functionality. Here's a screenshot of Claude estimating calories for some fries I photographed and logging them to MyFitnessPal automatically:

Screenshot of using MyFitnessPal integration in Claude
Screenshot of using MyFitnessPal integration in Claude

In claude.ai, I've created a project and been using memory to teach Claude about my diet, which is fortunately quite repetitive.

Check out the docker-compose setup and MyFitnessPal MCP server fork on GitHub. The MyFitnessPal MCP server is included as a submodule.

Edward J. SchwartzComputer Security Researcher7 min. read

New open LLMs are released constantly and keep improving. Gemma 4 was recently claimed to be groundbreaking, but when I tried the quantized version for coding agents like opencode, it was completely unusable—it gets stuck in output loops or unable to call tools with correct syntax. This is the norm, not an exception. Most quantized open LLMs I try for agentic AI simply don't work. Finding a setup that does requires trial and error across model, quantization, and dozens of settings.

I just want a command I can run to get a working LLM for my GPU. No hours of experimentation. No guessing at combinations. Just a proven setup. I couldn't find one, so I built auto-bench.

auto-bench

Auto-bench is a tool that allows you to define experiments, automatically run LLM inference servers with the proper settings, and execute a set of benchmarks against them. Rather than reinventing the wheel, I'm currently using Harbor Framework to run the tests. Auto-bench has first-class support for quantized models. This is important, because most existing benchmarks and leaderboards don't consider quantization, even though that is how many people run models.

My project is in its earliest stages, but I have at least one experiment to share: testing various quantizations of the Qwen3.5-2B model on a single problem instance from SWE-bench Verified (swe-bench/sympy__sympy-22914). I deliberately selected an easy instance to see if quantized models can perform basic tool calls to solve an easy problem.

Here is how this experiment is configured in auto-bench:

# Benchmark 22 quants of Qwen3.5-2B on a single SWE-bench instance
# Usage: auto-bench run configs/qwen-2b-quant-sweep.yaml

name: qwen-2b-quant-sweep
backend_type: llamacpp
dataset: SWE-bench/SWE-bench_Verified

instance_ids:
  - swe-bench/sympy__sympy-22914

model:
  name: Qwen3.5-2B
  source: huggingface
  repo_id: unsloth/Qwen3.5-2B-GGUF
  sweep:
    - label: BF16
      filename: Qwen3.5-2B-BF16.gguf
    - label: Q3_K_S
      filename: Qwen3.5-2B-Q3_K_S.gguf
    - label: Q5_K_M
      filename: Qwen3.5-2B-Q5_K_M.gguf
    - label: Q5_K_S
      filename: Qwen3.5-2B-Q5_K_S.gguf
    - label: Q6_K
      filename: Qwen3.5-2B-Q6_K.gguf
    # ... 17 more quantizations

sampling:
  temperature: 0.7
  top_p: 0.8
  top_k: 20
  min_p: 0.0
  presence_penalty: 1.5
  repetition_penalty: 1.0
agent:
  agent: openhands
  env: docker
  attempts: 4
  limit: 1
  setup_multiplier: 10.0

evaluation:
  run_evaluation: true

Part of my goal is to include all information needed to actually run the models properly. For example, the sampling section includes the sampling parameters that are recommended by Qwen for best performance, and these types of details can make a huge effect! My vision is to eventually have a leaderboard that will provide you with a llama.cpp command-line to run the model with the proper settings, and then you can just copy and paste that command to get a working LLM for your coding agent.

Results

Before diving into the data, here's what the columns mean:

  • Resolved: Number of problem instances successfully resolved by the agent
  • Total: Total number of attempts (8 in this case)
  • % Resolved: Resolution rate as a percentage
  • PPL (Perplexity): Measures the model's confidence in its predictions. Lower is generally better, though surprisingly this doesn't always correlate with task success
  • KL: Kullback-Leibler divergence—how much the quantized model's output distribution diverges from the original model's. Lower is better, but as we'll see, it's not a strong predictor of task performance
  • Runtime: Total time to run all attempts
  • Exceptions: Types of errors encountered (e.g., timeouts, exit code errors)
QuantResolved%PPLKLRuntimeExceptions
BF160/80%13.383m 37s
IQ4_NL1/812.5%13.670.03093m 53s
IQ4_XS0/80%13.680.031851m 53sTimeout, ExitCode
Q3_K_M2/825%14.330.077451m 48sTimeout
Q3_K_S5/862.5%15.080.133451m 39sTimeout
Q4_04/850%13.910.045419m 5s
Q4_13/837.5%13.680.027311m 14s
Q4_K_M1/812.5%13.790.02304m 8sExitCode
Q4_K_S2/825%13.780.02744m 43s
Q5_K_M8/8100%13.460.008251m 48sTimeout
Q5_K_S6/875%13.490.01008m 53s
Q6_K7/887.5%13.480.003551m 54sTimeout
Q8_04/850%13.390.001251m 49sTimeout
UD-IQ2_M0/80%17.610.267751m 56sTimeout(5)
UD-IQ2_XXS0/80%27.110.701851m 58sTimeout(6), ExitCode
UD-IQ3_XXS0/80%15.310.154951m 49sTimeout
UD-Q2_K_XL0/80%17.150.238851m 49sTimeout(2)
UD-Q3_K_XL6/875%13.940.052051m 56sTimeout(2)
UD-Q4_K_XL6/875%13.600.01647m 40s
UD-Q5_K_XL8/8100%13.510.007718m 2s
UD-Q6_K_XL3/837.5%13.480.00206m 51s
UD-Q8_K_XL3/837.5%13.370.00115m 49s
vllm1/812.5%6m 32s

The Base Model Is Broken

The most striking finding is that the unquantized base model (shown as vllm in the results) achieves only 12.5% resolution—worse than most quantized versions. BF16, which is nearly the original model without quantization, also consistently fails at 0%. This suggests the base model is fundamentally broken for this coding task, but quantization somehow fixes it.

Many medium-sized quantizations (Q5_K_M, UD-Q5_K_XL, Q6_K) achieve 100%, 100%, and 87.5% resolution respectively. Yet larger quantizations like Q8_0 fail again at 50%. This isn't about bigger being better—it's about finding the quantization that repairs the base model's broken reasoning.

KL Divergence Doesn't Predict Success

The KL (Kullback-Leibler) divergence column measures how much a quantized model's output distribution diverges from the original. If the base model is broken for this task, then staying close to the original (low divergence) just means inheriting the same brokenness. That could explain why there's no strong correlation between divergence and success.

Q5_K_M achieves 100% resolution with low divergence (0.0082)—but so does UD-Q5_K_XL with similar low divergence. Meanwhile, BF16 (essentially 0 divergence, nearly the original) fails completely at 0%. Some high-divergence models like UD-IQ2_XXS fail too, but others like Q3_K_S achieve 62.5% with divergence of 0.1334.

KL Divergence vs number of resolved attempts
KL Divergence vs number of resolved attempts

The plot tells the story: failing models (0 resolved) scatter across the entire divergence range—some very close to the original, some far away. If you stayed loyal to a broken base model, you'd fail. If you accidentally diverged in the right way, you'd succeed. KL divergence alone can't tell you which happened.

Multiple Failure Modes

There are two distinct failure modes visible in the results.

Failure Mode 1: Infinite Loops (AgentTimeoutError)

Many quantizations exhibit infinite looping behavior, where the agent gets stuck generating the same outputs repeatedly and eventually hits the timeout limit. Models like UD-IQ2_M, UD-IQ2_XXS, IQ4_XS, and IQ3_XXS show multiple AgentTimeoutError instances. Interestingly, this failure mode appears to correlate strongly with extremely aggressive quantization (e.g., IQ2 variants with very high KL divergence > 0.26).

Failure Mode 2: Silent Failure

The second failure mode is when the agent runs to completion without timing out, but simply fails to correctly solve the problem. Models like BF16, UD-IQ2_XXS, and UD-IQ3_XXS never produce output loops, but they still achieve 0% resolution. This suggests that the quantization has degraded the model's reasoning ability below a critical threshold where it can't effectively reason about code, even if it's still syntactically generating valid tool calls.

Conclusion: The Base Model Is Broken, Quantization Fixes It

The core finding is that the unquantized base model (12.5% resolution) and near-original BF16 (0% resolution) both fail for this coding task. Yet specific quantizations like Q5_K_M and UD-Q5_K_XL achieve 100%. Quantization isn't degrading a working model—it's repairing a broken one.

Notice in the visualization: models that fail (0 resolved) scatter across the entire KL divergence range, from very close to the original all the way to extremely divergent. Models that succeed tend to cluster at low divergence. But the scatter on the left side proves you can't predict failure from divergence—some quantizations stay very close to the original yet still fail.

The lesson is that model quality for coding agents is challenging to predict. This is why auto-bench exists—to empirically measure what actually works for your specific use case.

Open Questions and Limitations

This experiment demonstrates an interesting phenomenon, but it's based on a single problem instance from a single model family. The findings should be interpreted with appropriate caution:

  • Generalization to other models: Do these patterns hold for Llama, Mistral, or other model families? The behaviors might be Qwen-specific.
  • Generalization to other instances: I deliberately chose an easy instance to see if quantized models could work at all. Would the patterns hold on harder instances? SWE-bench Verified spans easy to extremely difficult problems.
  • Generalization to other tasks: Would we see similar results on SWE-bench instances beyond Verified, or on other benchmarks like HumanEval or MBPP?
  • Sampling parameter sensitivity: How much of the improvement from quantized models comes from properly-tuned sampling parameters? A controlled ablation would be valuable.

I'm actively running experiments to answer these questions! The auto-bench framework is designed to scale to hundreds of model/quantization combinations and thousands of problem instances. Stay tuned for results on larger model families, more problem instances, and more task types.

I'm happy to announce that my student Luke Dramko's paper "Idioms: A Simple and Effective Framework for Turbo-Charging Local Neural Decompilation with Well-Defined Types" has been accepted to NDSS 2026! Put simply, the paper shows that neural decompilers benefit greatly from explicitly predicting and recovering user-defined types (structs, unions, etc.) referenced in decompiled code.

Paper & Code

The paper has a great motivating example that I'll borrow here. The example starts with a C function that uses a struct type:

struct hash {
  int hash_size;
  int item_cnt;
  struct gap_array *data;
  int (*hash_make_key)(void *item);
  int (*cmp_item)(void *item1, void *item2);
};
struct gap_array {
  int len;
  void **array;
};
  
int hash_find_index(struct hash *h, void *item) {
    void *cnx;
    int index = hash_make_key(h, item);
    int cnt = 0;
    cnx = gap_get(h->data, index);
    while (cnx != NULL) {
        if (cnt++ > h->hash_size) return -1;
        if (!h->cmp_item(cnx, item)) break;
        index = hash_next_index(h, index);
        cnx = gap_get(h->data, index);
    }
    if (cnx == NULL) return -1;
    return index;
}

If you've read this blog, it will probably not surprise you that even state-of-the-art decompilers like Hex-Rays struggle to recover this code in a human-readable form. Here is the output from Hex-Rays:

__int64 __fastcall func4(__int64 a1, __int64 a2) {
  int v2;          // eax
  __int64 result;  // rax
  int v4;          // [rsp+10h] [rbp-10h]
  unsigned int v5; // [rsp+14h] [rbp-Ch]
  __int64 i;       // [rsp+18h] [rbp-8h]
  v5 = func2(a1, a2);
  v4 = 0;
  for (i = func1(*(_QWORD *)(a1 + 8), v5); i;
       i = func1(*(_QWORD *)(a1 + 8), v5)) {
    v2 = v4++;
    if (v2 > *(_DWORD *)a1) return 0xFFFFFFFFLL;
    if (!(*(unsigned int(__fastcall **)(__int64, __int64))(a1 + 24))(i, a2))
      break;
    v5 = func3((_DWORD *)a1, v5);
  }
  if (i)
    result = v5;
  else
    result = 0xFFFFFFFFLL;
  return result;
}

There are a lot of problems with this output, but the most serious is that the information about the hash struct has been completely lost.

A very exciting line of decompilation research is neural decompilation, which leverages neural models to either (1) directly decompile code, or (2) improve the decompiled code from traditional (non-neural) decompilers. I am personally extremely excited about the latter approach, which uses neural models to post-process the output of traditional decompilers such as Hex-Rays. Traditional decompilers have been studied for decades, so why not leverage their strengths while using neural models to fix their weaknesses? One popular example is the LLM4Decompile models. Here is an example of LLM4Decompile's output for this function:

int FUN_00100155(struct FUN_0009ff84 *VAR_0,void *VAR_1){
  int VAR_2;
  int VAR_3;
  void *VAR_4;
  VAR_2 = FUN_0009ff86(VAR_0, VAR_1);
  VAR_3 = 0;
  VAR_4 = FUN_0009ff88(VAR_0->VAR_5, VAR_2);
  while (VAR_4) {
    if (VAR_0->VAR_6 < VAR_3) { return -1; }
    if (!VAR_0->VAR_7(VAR_4, VAR_1)) { break; }
    VAR_2 = FUN_0009ff89(VAR_0, VAR_2);
    VAR_4 = FUN_0009ff88(VAR_0->VAR_5, VAR_2);
    VAR_3++;
  }
  if (VAR_4) { return VAR_2; }
  return -1;
}

Unlike Hex-Rays, LLM4Decompile correctly identifies that the function's arguments are pointers and that the first argument is a pointer to a struct. But what is struct FUN_0009ff84 and what are its fields, VAR_5, VAR_6, and VAR_7? And perhaps most importantly for reverse engineers, what offsets are those fields at? This information is crucial for understanding the code, but it has been omitted by the model.

One of Idioms' main contributions is to modify the training process of neural models so that they produce well-defined types such as structs with named fields. Unsurprisingly, this makes decompiled code much easier to understand. Here is the output of Idioms on this function:

struct hash_t {
  int size;
  int count;
  struct hash_table_t *table;
  int (*hash)(void *key);
  int (*cmp)(void *key1, void *key2);
};
struct hash_table_t {
  int size;
  void **items;
};
int hash_find(struct hash_t *hash, void *key) {
  int index = hash_index(hash, key);
  int i = 0;
  void *item = hash_get(hash->table, index);
  while (item != ((void *)0)) {
    if (i++ > hash->size) { return -1; }
    if (hash->cmp(item, key) == 0) { break; }
    index = hash_next(hash, index);
    item = hash_get(hash->table, index);
  }
  return (item == ((void *)0)) ? -1 : index;
}

An Unexpected Advantage of Joint Predictions

Perhaps surprisingly, in addition to improving the readability of the code, jointly predicting both code and types also significantly improves the accuracy of the decompiled code!

Across multiple models and evaluation metrics, Idioms consistently outperforms prior neural decompilers:

  • On ExeBench: Idioms achieves 54.4% test-pass accuracy (vs. 46.3% for LLM4Decompile and 37.5% for Nova).
  • On RealType: A dataset we introduce that contains substantially more and more realistic user-defined types (UDTs); Idioms improves correctness metrics by 95–205% over prior work.
  • Context helps: Adding neighboring-function context improves UDT recovery—up to 63% improvement in structural accuracy—with little downside for larger models.

Beyond Decompilation

Surprisingly (to me), Idioms also outperforms standalone type recovery tools such as Retypd, BinSub, TRex, and TypeForge, by at least 73%. This suggests that generative, context-aware approaches may be well suited to resolving the inherent ambiguity of type recovery than prior approaches, even though this was not the original motivation of Idioms.

Edward J. SchwartzComputer Security Researcher3 min. read

(Apologies in advance, this is probably going to be rambling.)

I spend a lot of time looking at various research artifacts. Yesterday, I was looking at several verification artifacts, and I was struck by how much impact small details like a Docker container can have. One of the projects I was using is STOKE. STOKE is a cool project, but the salient detail for this post is that it is an abandoned research project. The last commit was in December 2020. This is very common: Ph.D. student creates a project, maintains project, graduates, and then no longer maintains project. Despite that, STOKE has a Docker container, which makes using the project trivial.

In contrast, I was also attempting to run Psyche-C this week. Like STOKE, there's a fair bit of bitrot on the branch that contains the type-inference component. Unlike STOKE, Psyche-C does not have a Docker container. Part of this branch uses Haskell lts-5.1, which is from 2016! Trying to get this running was a nightmare, since modern versions of stack, GHC and cabal could not cope with such an old environment. I was eventually able to get it running by creating an Ubuntu focal docker image but it took me an entire day. I also created a HuggingFace space for it.

I have said it before, but I just love HuggingFace spaces for hosting research artifacts. It makes it almost effortless for others to try out your research. I wish more researchers would use them.

I think that the decompilation and reverse engineering research community could also significantly benefit from using HuggingFace spaces and generally making artifacts easier to use. I say this because there are many subtle details about reverse engineering research artifacts that can make them less usable in practice.

For example, I was recently reading DecompileBench, which is a good paper about benchmarking decompilers. In particular, they have a very clever method for testing whether a decompiled function is semantically equivalent to the original source code. In short, they compile the decompiled function in isolation and splice it into the original program, and do some testing to see if it behaves the same way. I've been thinking about this topic a lot recently, since I have been talking with some of my students about it. The problem is that if the binary is stripped, the decompiler can't refer to symbols by their original name, and thus the decompiled code can't reliably be linked back into the original program. (Ryan pointed out on Bluesky that this is possible in some cases.) DecompileBench ignores this problem and decompiled unstripped binaries. This is a problem, because decompilers are usually used on stripped binaries, and they generally perform significantly worse on them.

My goal is not to criticize DecompileBench; I think it's a nice paper. My point is that there are many subtle details like this that can make research artifacts less useful in practice. I've had my own share. As one example, the DIRT dataset was stripped using the wrong command, so that function names were still present in the binary, which is unrealistic. Fortunately, it turned out (surprisingly) that this did not significantly affect the results in that paper, but it could have.

I think part of the problem with these two examples is that it's hard to get close to the actual use case with these projects. In decompilation, the real use case is decompiling stripped binaries in the wild. But it's hard to run DIRTY on a new binary to see how well it works. I have found this to be a common problem in machine learning-based research. The straight-forward approach is to start with a dataset for which you have ground truth, and then train and test on that dataset. This often leads to preprocessing code that expects to have the ground truth information available. This is problematic when you want to perform inference on a new example when you don't have ground truth information, e.g., the primary use case of these technologies!

My overall message is that docker containers and HuggingFace spaces are great ways to make research artifacts easier to use. This is important in general, but it's also important to be able to get as close to the real use case as possible. If, for example, your technique only works for unstripped binaries and you forget to mention this in your paper, a docker container or space is going to make that very apparent.

I have a pretty hot take: the top-tier conferences should mandate that research artifacts be easy to use, e.g., via docker containers or HuggingFace spaces, on new examples, and that these artifacts should be considered as part of the submission. Having a separate, optional artifact evaluation process simply doesn't work. (The incentive for going through the artifact evaluation is a badge, which is essentially a sticker for grown-ups!) But if reviewers can actually try out the artifact on new examples, they can see how well it works in practice. This would significantly improve the quality of research artifacts in our community.

Powered with by Gatsby 5.0