← Blog

Physical Commodities Trade Allocation as an Optimisation Game

Physical Commodity Trading Allocation as an Optimisation Game

The PnL of a physical trading desk is directly influenced by the quality of its allocation decisions, which is essentially about finding the best match for your purchased tonnes (your supply) to your sales tonnes (customer demand) for max PnL.

This naturally lends itself to an optimisation framework. Compared to trying to model prices in the paper world, the physical allocation problem is much more bounded and structured, as it is directly anchored by real assets and real-world constraints.

Once you enter into a physical trade, you are tied to an actual cargo with a specific location, quality specification, counterparty, delivery window, QP, etc. From there, a physical trader should already have a rough sense of the eventual PnL as the available options are constrained by physical factors such as freight, financing, timing, storage, credit, etc. and the paper side is typically hedged out.

Thus, it makes a lot of sense that if trading desk information can be structured in a way that AI can understand, AI can help traders evaluate a much larger number of allocation possibilities very rapidly, and potentially identify opportunities outside of what the trader initially saw.

Optimisation Game, not Problem ;)

The inspiration for this solidified after I created a gamified version of physical commodities trading in my TradingDesk-Sim post.

The original game has been since revamped to be a better mirror to an actual physical desk, to play out the market scenario in 2025 where we had very exciting CME/LME arb opportunities.

You can try to play the human playable version of the game there, just be warned it might feel like work haha

For linguistic consistency I will stick to calling this task the allocation game rather than allocation problem

Trying out Local Models (and pulling my hair out)

I thought my experience benchmarking models for SnD updating performance would have made it much quicker to get AI to play this allocation game, but it ended up taking far longer than expected. This was largely due to my stubbornness in wanting to get local models to play.

It was super frustrating that the local models (Qwen, Bonsai) I tried to run on my mac-mini just kept failing, even when ChatGPT models were able to play with no issues.

A lot of effort went into infrastructure and model interface design to try accommodate local model shortfalls, including creating a JSON-based version to strip away any interface-related challenges, but even after significant infrastructure changes, my local models could not make it to the end of the game before crashing.

The problems encountered in this process that stood out to me were:

  1. My local models were very easily overwhelmed. A trading decision requires a lot of information e.g. quantities, grades, locations, freight, deadlines, QP, and as the game progressed, the inputs became too big as more and more previous decisions had to be remembered.

    E.g. a Qwen run had a 43,416 token prompt truncated to 20,482 tokens, which meant it simply did not receive the full instructions required to make a sound trading decision.

    To work around that, I explored using compact records, paginated inspections and checks that reject oversized inputs rather than silently truncate them.

    I also experimented with using a 131,072-token YaRN Context Extension, which seemed to have solved the context capacity issue. However, that did not help prevent Qwen from making basic allocation errors, e.g. trying to reuse a cargo that an earlier order had already consumed.

  2. They also really liked to talk rather than actually execute decisions. The local models spent significant inference time (62% of Qwen’s, 45% of Bonsai’s) making summaries of what they did instead of making trading decisions.

    Actually, they like to talk so much that they can explain what they intend to do, but end up giving an invalid or different command. And they also made up things, e.g. Qwen invented cargo IDs.

Why Local Models

Why was I so stubborn on making local models work, you ask? The reason is that this task requires access to a significant amount of commercially sensitive information which should ideally never leave the firm (e.g. contracted volumes, QP, premium, counterparties). It just felt very uneasy to me to be sending such critical data to OpenAI servers, as much as they say they do not use your data for enterprise plans.

If local models can do the job, there would be no need for any data to be sent to someone else’s servers. This goes back to my explorations around data architecture for commodities firms

More critically, local models also have the advantage where you have full visibility of and can tweak its model weights, i.e. you unlock the ability to train the model for your specific purposes, and in doing so that model can become your edge.

For now, the local models that I can run on my mac-mini cannot seem to do this, and I think it likely stems from the fact that LLMs are trained to make useful looking text rather than actually making decisions autonomously.

It is absolutely possible though, that much like how AlphaGo is a trained model to play the game of Go, there can be a “AlphaTrade” trained to play the game of physical commodities allocation. The point of local models is that they can be trained and tweaked for specific tasks.

Stay tuned for benchmarking on the close weight models…

Right now GPT-6 Astra can already play the allocation game till the end. I will test more models soon, benchmarking post to follow!