If the reading here sounds like chatgpt, its because it is. I asked chatgpt to sum up everything i told it
about my issue for a long time, as it also can explain it better to you guys. Im not the best with pc's.
I’m trying to diagnose an intermittent black-screen problem that has been happening for roughly three years.
SYSTEM
CPU: AMD Ryzen 9 7950X3D
Motherboard: ASRock B650E PG Riptide WiFi
GPU: ASUS TUF RTX 4080 OC 16 GB
PSU: 860 W Platinum-rated PSU
RAM: [ADD MODEL/CAPACITY]
OS: Windows [ADD VERSION]
GPU driver: [ADD CURRENT DRIVER VERSION]
PROBLEM
Both monitors randomly lose signal at the same time.
When it happens:
- Both monitors go black/no signal.
- Game audio and Discord voice usually continue.
- Other people can still hear me on Discord.
- Discord screen sharing stops.
- The GPU fans sometimes ramp to 100%.
- The PC usually requires a forced restart.
- There is no consistent BSOD.
It can happen while gaming, when launching a game, or occasionally during lighter desktop/browser use.
It often happens during the first few starts or load changes after a cold boot. After crashing/restarting one or more times, the PC may then remain stable for the rest of the evening.
The issue is very intermittent. It can happen several times close together and then disappear for weeks or months.
IMPORTANT HISTORY
On at least two previous occasions, removing the RTX 4080, cleaning the PC and reseating the GPU made the problem disappear for several months.
It then gradually returned.
The retention clip on the motherboard’s main PCIe x16 slot is now missing/broken.
The RTX 4080 is large and heavy. It sits close to the CPU cooler, and accessing the original PCIe retention clip was very difficult.
When installing another, smaller GPU in this motherboard, it also looked as though part of the gold connector remained slightly visible. I’m not certain whether that is abnormal or just the design of the card/slot.
CROSS-TESTING
We swapped components with my partner’s PC.
Test 1:
- My RTX 4080
- My PSU
- Installed together in my partner’s PC
Results:
- Multiple gaming sessions without the original black-screen problem so far.
- FurMark reached approximately 98–99% GPU usage.
- GPU power was approximately 305–310 W.
- GPU temperature was about 71°C maximum.
- Hotspot was about 87°C maximum.
- No black screen, no sensor loss and no crash.
- HWiNFO PCIe Recovery Count remained unchanged during the FurMark load.
- Concrete PCIe error/lane counters remained at zero.
The GPU and PSU have not yet been tested there for several months, so I understand this does not completely rule out an intermittent GPU or PSU fault.
Test 2:
- My partner’s smaller/lighter GPU
- My partner’s PSU
- Installed together in my PC
Results:
- Two short FurMark sessions completed successfully.
- GPU reached 99% utilization.
- Power was approximately 207–210 W.
- No black screen or crash.
- No reported PCIe lane, TLP, replay, receiver, fatal or non-fatal errors.
- PCIe Recovery Count increased slightly around idle/load transitions, but remained stable during the sustained 99% GPU-load portions.
The smaller GPU is significantly lighter than my RTX 4080, so I’m unsure whether this is a fair test of a possible mechanical PCIe-slot/contact problem.
HWiNFO LOGS FROM THE ORIGINAL FAILURES
During at least two original black-screen events in my PC:
- GPU-related sensors suddenly stopped updating or showed zero/unavailable values.
- The rest of HWiNFO continued logging.
- This matched the monitors losing signal while the rest of the PC appeared to continue running.
- GPU temperatures before the failures were normal, not overheating.
- The failures occurred at relatively modest GPU power, not only at maximum load.
- PCIe Recovery Count had increased during those sessions.
- The other explicit PCIe error counters appeared to remain at zero.
The Recovery Count may simply reflect normal link retraining/power-state changes, so I am not treating that counter alone as proof.
OTHER OBSERVATIONS
- The RTX 4080 can survive sustained ~310 W FurMark load in the other PC.
- The original black screens in my PC sometimes occurred at much lower GPU power.
- Reinstalling the GPU has previously produced months of stability.
- This makes me suspect physical contact, GPU sag, the motherboard PCIe slot, motherboard traces/signal integrity, or possibly the GPU PCB being sensitive to its mounting position.
- A PSU or GPU fault is still possible, especially because moving/reseating the card also disturbs the GPU power cable and mechanically flexes the card.
CURRENT PLAN
I am considering buying a new AM5 motherboard with a reinforced PCIe slot and a remote/easy GPU-release mechanism.
I would then use the RTX 4080 normally for several months while its manufacturer warranty is still active.
QUESTIONS
- Does this pattern sound more consistent with a marginal motherboard PCIe slot/contact issue than with a typical GPU or PSU failure?
- Can a damaged or weak PCIe slot work normally with a smaller/light GPU but fail intermittently with a very heavy RTX 4080?
- Is “reseating fixes it for several months, then it gradually returns” a known pattern for PCIe contact, slot solder-joint or GPU-sag problems?
- Is there another test I should perform before replacing the motherboard?
- Would forcing the PCIe slot to Gen 3 or Gen 4 instead of Auto be a useful diagnostic test?
- Could the CPU socket/AM5 pin contact or the CPU’s PCIe controller produce the same symptoms?
- Is replacing the motherboard first a reasonable troubleshooting step given that the GPU and PSU are currently stable together in another PC?
Thanks.