I only care about one single approach, which is the one I’d be going for: runtime generated code running on a real-time priority process owning a whole core.
- Allows for easy in-order execution, honouring the timestamps.
- Removes most, if not all, memory reads.
- Would allow late insertion of an “event” at the cost of thrashing the cache. Can not actually recommended this.
- Can easily be mixed with Python code, if required. No, I’m not talking about runtime generated python code. Talking strictly native x86.
From there I’d just expand.
More cores.
More nodes.
I’ve been doing that “everywhere I came across” at least once.
Linux and Windows aka x86 and on ARM, specifically on the Nintendo DS.