01 / Question

Why might two useful adapters fail when combined?

LoRA lets us fine-tune a model for a task without retraining all of its weights from scratch. The result is a small adapter that can be stored, shared, and applied to the same base model later. As agents acquire more skills, this makes adapters an appealing way to add capabilities without keeping a separate full model for each one.

Recent work has shown that adapters can be combined to manage growing agent skill sets. We wanted to look at a stricter version of that idea: what happens when adapters are composed dynamically at test time to tackle a task that involves functionally composite transformations?

02 / Hypothesis

The outcome may depend on how the two skills meet.

A weighted merge combines parameter updates, but it does not create a visible handoff from one operation to the next. The result may depend on how much the updates overlap, which parts of the model they change, how the two tasks alter the record, and whether one skill dominates the other after merging.

Our hypothesis was that these factors can help explain whether two combined adapters perform well on functionally composite tasks. We expected some merges to preserve both atomic skills, some to favor one, and some to lose useful behavior altogether. A directly trained composite adapter gives us the reference point: if it learns the target, a dynamic failure is evidence about the combination rather than only about the task.

We are interested in patterns in task relationships, retention, interference, and adapter geometry that separate one kind of failure from another.

03 / Experimental setup

A small task world where every answer can be checked.

To test that hypothesis, we evaluated adapters on deterministic key-value records. These records give the project a shared language for the tasks and make the intended intermediate and final answers inspectable. The atomic tasks include sorting fields, selecting values, lowercasing values, removing duplicates, renaming keys, and swapping keys with values.

We then paired those operations to create composite tasks. For example, the selection-then-sorting task asks the model to keep vowel-initial values and sort the survivors. Its generator requires at least two selected fields and makes sure sorting changes their order, so neither step is incidental.

We used Qwen2.5-0.5B-Instruct with rank-8 LoRA updates on attention projections. Independently trained adapters were combined with additive and CAT methods at several weights. Dedicated composite adapters and joint controls were trained on the combined examples so we could tell a merge problem from a task or training problem.

We scored structured content as well as exact strings, then saved atomic retention, interference, parsing, endpoint, and update-geometry diagnostics with the predictions.

Record inspectorselection + sorting
promptFirst select vowel-initial values, then sort the selected fields.
input record
a:echob:alphac:hoteld:bravo
selected intermediate
a:echob:alpha
dynamic merge
a: echob: alpha
selection retained; order missed
dedicated composite
b: alphaa: echo
ordered target
target contentb: alpha · a: echo
Try it Choose an atomic prompt or the ordered composite. The input record moves through the relevant operation, then the two result cards show the difference between the adapters. A dynamic composite is made by merging the independently trained adapters at test time, so it can retain a fragment of a skill without producing the ordered target. A dedicated composite is trained directly on examples of both operations and provides the learnability control. The records here are illustrative, not individual model predictions.
04 / Results

The controls learned the task; the dynamic merges didn't.

With those controls in place, the result was clear across the five retained task-pair families. Additive and CAT dynamic adapters scored 0% on the composite tests, while adapters trained directly on the combined examples reached roughly 98-100% on IID data. The transformations and instructions were learnable. The problem appeared when the two independent updates had to cooperate at inference time.

The failures also had different shapes. For sorting followed by selection, the merged model kept about 81% of its sorting performance but only about half of its selection performance. Reversing the order changed which skill was weaker, while the composite score stayed at 0% in both directions.

The atomic controls showed whether a merge erased both skills, leaned toward one, or retained pieces of each without producing the intermediate record the final task requires.

05 / What surprised us

The failures were uneven, which made them more informative.

The first headline is easy to summarize: the dynamic composite failed. The more useful observation came from looking underneath it. One behavior could remain usable while the other weakened, and the model could still miss the task that asked them to interact.

That unevenness rules out a simple story in which merging either preserves everything or destroys everything. It also gives us a more concrete way to think about interference. If sorting survives at roughly 80% while selection falls to roughly 50%, the merged model has not become useless. It has become biased toward one part of the requested program.

This changes what a useful composition report should contain. The composite score says whether the final answer was right. Atomic prompts show what the model can still do, retention differences show which adapter the merge has favored, and the saved predictions let us inspect the exact point where the final answer diverges.

06 / Limits and next steps

The result is useful, but the world is small.

There are clear boundaries around what we have learned. The experiments use one small instruction model, one LoRA setup, synthetic records, and two static merge methods. The task pairs reuse some of the same adapters, so they are not five independent tests. The geometry measurements also cover attention projections only.

Those limits suggest a natural next comparison rather than a long list of extra sweeps. Keep the records and evaluation fixed, then place static merging beside explicit chaining and task routing. Matched runs with MLP-targeted adapters, larger models, and more task families would tell us which parts of the pattern persist.

For now, the narrow conclusion is that retaining capabilities and composing them are separate properties. The next step would be to find out whether an explicit execution mechanism can preserve the intermediate state that a static merge seems to lose.

07 / Code and paper

Code and paper

The project includes the task generators, adapter weights, composition runs, predictions, structured evaluation, qualification reports, and update analysis. The packaged data makes it possible to inspect the examples behind the aggregate results and see how a particular merge failed.

Code Paper Data and runs Related field note