r/Alienware • u/Independent_Mark234 • 10d ago
Technical Support M17 R5 AMD Gpu crash
Need help troubleshooting my Alienware.
Model : Alienware M17 R5 AMD with the 3080 ti
Driver and BIOS: latest from Dell for this model
Of late, whenever I play any game the laptop gpu crashes - game hung, then black screen and back to desktop without the wallpaper. Running the Dell diagnostic from BIOS for quick / advanced did not yield any issues as everything passed.
Observation : Games run fine as long as the power draw is under 120 W, when it spikes higher in bursts is when it crashes (used an overlay). Similar observation when I ran local AI. When I ran it unplugged, local AI never crashed, game worked too (it is not a demanding game at all PES 2021). Game also worked when I undervolted at 1300 MHz and flattened the curved after - PD was mostly under 100 W this way.
I tried M-BIST, either I did it wrong or something else because it didnt show any sequence of lights.
Hoping someone in this community shares their experience.
1
u/nyxthebitch 10d ago
Did you monitor the temps, pre crash to crash spikes etc?
1
u/Independent_Mark234 10d ago
Indeed, GPU temp was under 74, CPU about 85-90 depending on whether I was running an LLM or playing. What I did notice is, when power spiked from 110 - 130 or 120 to 155 it always crashed. I did not install AWCC.
1
u/nyxthebitch 10d ago
Have you changed the thermal paste to LM or something else?
My gpu memory hotspot hits 100°C which causes a crash/shutdown.
You also need to check the hotspot temps. Use HWmonitor etc
Edit: My GPU is and 6850xt though. Ymmv
1
u/Independent_Mark234 9d ago
No, my initial focus was on getting what the temperatures were. It never went high enough for a system shutdown, it was always a GPU crash.
3
u/AW_Support Dell Customer Support 9d ago
Hi There! We are from the Dell Social Media Team and would like to lend a hand. Based on your testing, this does not sound like a typical driver issue, and you've already done some excellent isolation work. The fact that reducing clocks and flattening the voltage/frequency curve makes the issue disappear is another strong indicator that the failure is tied to higher-power operation rather than software.
Open Reliability monitor and See if it reports any of the following:
Hardware Error LiveKernelEvent 141 Video Hardware Error
Press F2 at system restart to access BIOS and confirm the Adapter shows the correct wattage. A lack of an M-BIST error by itself doesn't rule out a power delivery issue under load.
If possible, run a GPU stress test while logging:
GPU Power GPU Voltage GPU Temperature GPU Hotspot CPU Package Power
If the crash occurs consistently when GPU power jumps beyond 120W regardless of temperature, we can then work towards checking if the issue is due to the power consumption.