E.g. we can build a board around a MC68000 where we make it lock up forever in a bus cycle, waiting for a DTACK that doesn't arrive.
Some early microprocessors had clocked bus cycles without handshaking. They would put out an address on some address lines and signal some line together with a read/write indication, and then expect the transfer to be completed within some clock cycles. If nothing is attached to the address, they would read whatever values are on the bus, like maybe all 1's if it is an open drain system that requires the transmitting device to pull to ground to indicate zero.
I'd say that kind of thing belongs to a hall of shame; it requires software hacks to interface with anything that can't keep up with the prescribed bus cycle.
IIRC, early 68k Macs had processor upgrades hooked onto the 68k bus and did exactly this - an early boot driver run by the “actual” CPU coordinated handoff with the expansion CPU’s bootstrap ROM.
https://en.wikipedia.org/wiki/Metastability_(electronics)
Ultimately... failure modes must arise because naive gate design simulation models can't determine issues in a computationally feasible time frame.
The consequences of "fixing" CDC prone design flaws makes a processor many times slower (8 to 16 times slower on my dumb attempt), and develops weird alien design features very different from Von Neumann architectures.
This is why we can't have nice things. =3
> Trapped/emulated/virtualized instructions may only time the trap, not the handler.
But I feel like that 12ms write to an ACPI IO port at current leaderboard position 8 is probably trapping to SMM and being handled there.
Leads to x86 page table MMU magic being turing complete: https://github.com/jbangert/trapcc
And the simplest thing you can do on such a system is just to loop indefinitely, thus creating a simple instruction with a memory access (mov or anything, doesn't really matter, even the instruction fetch for a nop would work) to take infinite time.
Nope. Not on x86. You can use either physical or virtual addresses at your choosing. Consumer OSes use virtual ones, so you can swap out page tables (yes, really!). See https://wiki.osdev.org/X86_Paging "Page directory".
Not an LLM problem, just an undecaffeinated meat brain and some faulty memories. I've misread 'When PS=0, the page table address field represents the physical address of the page table that manages the four megabytes at that point.' to mean that when PS=1, the address isn't physical. But PS is page size... And I somehow remembered that you could induce pagefaults when walking the page tables...
Sorry.
I wonder what the actual limit on this `fxrstor64` is right now. If you can stall the PCIe bus for that long, then why not indefinitely? Certainly there's no forward progress guarantee here.
You're probably confusing with:
- POP CS being broken and later becoming a prefix
- some opcodes being "reserved NOPs", i.e. reserved without generating #UD. They are used for instructions that may be defined in the future while guaranteeing backwards compatibility, for example new kinds of prefetches. MPX bounds checking instructions were also encoded in reserved NOPs.
It's interesting that a lot of ALU operations occupy four opcodes (memory source/memory destination x byte/word) but 0x84/0x85 and 0x86/0x87 only need two because not only are they commutative, but also 0x84/0x85 do not write to any operands and 0x86/0x87 write to both. So there is no difference between memory as the source or destination operand.
The actual typical hardware implementation just fetches the next 16-32 bytes from icache, shifts it to the correct alignment, and slams it into a bunch of parallel decoders which each attempts to decode one x86 instruction per byte.
The next cycle, the first 1-6 non-overlapping valid instructions are accepted into a queue for further decoding. The NOP almost certainly takes up space in this queue.
At no point does RIP get incremented by one. There isn't even a single physical RIP register to increment, the CPU is "executing" dozens or even hundreds of RIPs in parallel.
There was a boy A very strange, enchanted boy They say he wandered very far Very far, over land and sea A little shy and sad of eye But very wise was he.
And then one day A magic day he passed my way And while we spoke of many things Fools and kings This he said to me: "The greatest thing you'll ever learn Is just to love and be loved in return."
Fools QEMU and makes for a good "am I on a VM" logic test.
A function of precalculating XOR operand inadvertly twice at QeMU TLB compute time AFTER retrieval of and toward its cached IMUL operand value.
In short, emulation doing preparation of registers twice (negating XOR)
Now you have a logic test revealing QEMU thru minute differential of IMUL operand value and its different multiplication results.
ROT13, anyone?
Disclaimer: works only on RXW memory page. It is literally a self-modifying code.
My guess... On Skylake, multiple in-flight RDTSC instructions slow each other down for some reason?
Possibly because it's attempting to provide a strict monotonic guarantee, that no two RSTSC instructions will return the same timestamp. Intel's manual only claims monotonic, which theoretically allows for two RSTSC instructions to return the same timestamp.
It would be much more interesting to know the results if you're only allowed to use main memory.
...and with things like https://en.wikipedia.org/wiki/ExpEther , you can get even higher latencies.
A very long-running instruction can be used to break SMI: https://github.com/xoreaxeaxeax/smiiiiiiiiiiiiiiii
What’s that law called about programmers wasting all the compute on abstraction?
There has been a ton of work to improve throughput, but that's often come at the cost of latency, because a great technique to improve throughput is batching, to avoid the per-task overhead cost.
We've also added many layers of abstraction.
A seminal example is Dan Luu's "computer latency" table (2017) https://danluu.com/input-lag/, which measures the time between a keypress and visible changes on the screen, and wherein an Apple IIe has latency 5x lower than a Lenovo X1 Carbon on Windows.
From some perspective, you could say that the Lenovo on Windows example is an impressive feat of engineering, considering all the subsystems involved (USB, interrupt dispatcher, input layer, windowing system, double-buffering...)
This throughput-over-latency tradeoff can be seen across the entire computer landscape design.
If I were designing an OS shell I would simply give the user the ability to easily make temporary text or image buffers and copy and paste content to/from them, rather than using a Notepad or a Paint.
If some app responds in 10ms or less, it is INTERACTIVE.
makes you think.
I looked it up and it is .1 seconds (100ms)
The basic advice regarding response times has been about the same for thirty years [Miller 1968; Card et al. 1991]:
- 0.1 second is about the limit for having the user feel that the system is reacting instantaneously, meaning that no special feedback is necessary except to display the result.
- 1.0 second is about the limit for the user's flow of thought to stay uninterrupted, even though the user will notice the delay. Normally, no special feedback is necessary during delays of more than 0.1 but less than 1.0 second, but the user does lose the feeling of operating directly on the data.
- 10 seconds is about the limit for keeping the user's attention focused on the dialogue. For longer delays, users will want to perform other tasks while waiting for the computer to finish, so they should be given feedback indicating when the computer expects to be done. Feedback during the delay is especially important if the response time is likely to be highly variable, since users will then not know what to expect.
from Jakob Nielsen:
https://www.nngroup.com/articles/response-times-3-important-...
less readable but the original paper:
https://www.yusufarslan.net/sites/yusufarslan.net/files/uplo...
Humans can perceive much smaller latencies.
If you look at the Card & Miller reference, at least some humans can perceive differences in ~50ms vs 100ms latencies when typing (in my limited testing, it’s likely you can!). There’s some newer research I don’t have handy that I believe found error rates decreased and NSAT improved until around at least 30ms (if not 20ms).
On that note, humans can definitely distinguish 60hz vs 120hz reliably (about 8ms faster per frame).
Even faster: with a reference (eg when dragging on a touchscreen), humans can distinguish down to at least 1ms vs 10ms of latency: https://m.youtube.com/watch?v=vOvQCPLkPt4
And you can probably distinguish metronomes that are off by about 1-2ms. Much smaller for other things (like metronomes that slightly slower or faster than one another).
This is a special interest of mine XD
How'd you get that number? USB defaults to polling at 125Hz and a lot of devices go at 1000Hz (or higher). The rest of the pipeline should be a fraction of a millisecond. I guess bad debouncing hardware can add a lot more, but that's far from USB's fault.