And in the comments, they note:
"I replaced the 18pF capacitors on one of the non-working boards with 10pF capacitors, but it still doesn’t boot or respond to the debugger. That surprised me – I thought that was the answer! It was a rush rework job and I made a bit of a mess of it, including accidentally desoldering and reinstalling the crystal, so I’ll try again later with another board. But it appears that the capacitor value may not have been the issue after all. Either 10pF is a bad value, or it’s the crystal itself that’s at fault, or I’ve failed somewhere in my troubleshooting reasoning. Hmm."
In my experience, mystery stability problems are often caused by capacitors, but of a different kind: decoupling capacitors on the power supply pins. If there's not enough of them to keep up with the noise originating from motors or digital switching, I'd expect that exact issue. Intermittent "impossible" CPU states on some boards, no rhyme or reason (because sometimes, that +/- 10% saves you and sometimes it does not). I'd try more caps and possibly some ferrite beads.
A multimeter can make it look like a chip is receiving the correct voltages, but there can be all sorts of noise and fluctuations that will cause really strange and inconsistent issues. You don't need a super fancy scope, even a 100mHz cheap handheld is more than good enough for checking power rails.
Also, for anyone who has to troubleshoot or fix boards, those smd resistor and capacitor kits are absolutely invaluable to have on the bench. After a decent scope, meter, and soldering station I would say that should be the next purchase for setting up an electronics lab. Nothing more annoying than trying to debug an issue and having to wait for the right value cap to be shipped. Especially since capacitors specifically often seem to have some trial and error to finding the right value for something like a decoupling cap. Fully modeling the noise generated on a supply rail often isn't completely possible, at least in my experience.
20 MHz is well within the range of any random scope, and you should have at least one "any random scope" on your desk if you're doing electronics seriously enough to reach the "diagnosing QA fail boards" stage. So probe the damn thing - and see if the crystal ever reaches a stable operating frequency once the board is powered. That would tell a lot.
And yes, the "holy book of SMD passives" is a must too. They cost you what, $40 each? And can save you days of waiting for the "right" passives to arrive when you need to test a circuit change. You don't need to have every size - footprints are negotiable when you're hand-soldering - but you should have at least one Book of Many Resistors and one Book of Many Capacitors.
But you'd need a few more than 100 millihertz for that lol
This is not a broken device. It didnt break. This is a batch of new devices that never worked to begin with straight from CM, and is being tested in standardized controlled rig all other freshly manufactured devices pass with no problem.
We had a new batch of hardware come in that had a failure rate of about 50%. Normally it was on the order of < 1%. Our CM was very good at troubleshooting.
Nothing had changed in the BOM so we were left scratching our heads.
Hooked up the JTAG debugger to see where it was failing to start, and the CPU wasn't even coming up. Power rails looked good, but the CPU just wasn't booting.
Eventually we discovered that the supplier had given us a batch of crystals that were slightly more sensitive to the capacitance, and our design was just on the margin of working.
Lowering the caps to the proper values according to the xtal's datasheet got everything working again.
1. Check the power rails are correct and stable (at the target devices, not the PSU source)
2. Check the resets are correct (polarity, level, sequencing) and reaching where they're needed
3. Check the clock is toggling cleanly (no jitter, has clean monotonic waveform)
4. General signal integrity and setup/hold of signals that are related to the issue
That catches a lot of basic issues. If those are all clean, then you go deeper.
To the article author's suspicion of crystals, I have seen crystal oscillators fail (stopping toggling) as ambient temperature ramps up and down; that's a nasty one to catch and prove, but it can happen. Changing vendor was the only solution there.
“Some exhibited “haunted” behavior, seemingly jumping to random sections of the mcu program code, outputting messages on the display that made no sense given the context. One of them appeared to work in slow motion, with LED blinking and display updates noticeably more sluggish than normal,”
my first thought was that it smelled like a clock issue.
Some of the nastier issues I have had the pleasure to debug included (a) traces that had microcracks which affected analog readings when the PCB heated up after prolonged usage (QC issue from the PCB fab) and (b) a (suspected) ESD strike that gradually took out several components in the weeks following as I was investigating the device while new problems kept popping up. Marginally stable composite amplifiers have also caused some headaches over the years.
Hardware really is hard.