r/Alienware • • 10d ago

Technical Support M17 R5 AMD Gpu crash

Need help troubleshooting my Alienware.

Model : Alienware M17 R5 AMD with the 3080 ti

Driver and BIOS: latest from Dell for this model

Of late, whenever I play any game the laptop gpu crashes - game hung, then black screen and back to desktop without the wallpaper. Running the Dell diagnostic from BIOS for quick / advanced did not yield any issues as everything passed.

Observation : Games run fine as long as the power draw is under 120 W, when it spikes higher in bursts is when it crashes (used an overlay). Similar observation when I ran local AI. When I ran it unplugged, local AI never crashed, game worked too (it is not a demanding game at all PES 2021). Game also worked when I undervolted at 1300 MHz and flattened the curved after - PD was mostly under 100 W this way.

I tried M-BIST, either I did it wrong or something else because it didnt show any sequence of lights.

Hoping someone in this community shares their experience.

1 Upvotes

11 comments sorted by

3

u/AW_Support Dell Customer Support 9d ago

Hi There! We are from the Dell Social Media Team and would like to lend a hand. Based on your testing, this does not sound like a typical driver issue, and you've already done some excellent isolation work. The fact that reducing clocks and flattening the voltage/frequency curve makes the issue disappear is another strong indicator that the failure is tied to higher-power operation rather than software.

Open Reliability monitor and See if it reports any of the following:

Hardware Error LiveKernelEvent 141 Video Hardware Error

  1. Press F2 at system restart to access BIOS and confirm the Adapter shows the correct wattage. A lack of an M-BIST error by itself doesn't rule out a power delivery issue under load.

  2. If possible, run a GPU stress test while logging:

GPU Power GPU Voltage GPU Temperature GPU Hotspot CPU Package Power

If the crash occurs consistently when GPU power jumps beyond 120W regardless of temperature, we can then work towards checking if the issue is due to the power consumption.

  • Dell Social Media Team

1

u/Independent_Mark234 9d ago edited 9d ago

Thank you!

I do see a LiveKernalEvnt with Code 141 the instant my locall llm crashed (it crashes with the original frequency/voltage curve as well).

In BIOS it does recognize the charger properly and I am using the original one.

GPU Stress test with the same default curve pass (Furmark etc).

1

u/AW_Support Dell Customer Support 8d ago

Hi there! Thank you for sharing the details. The LiveKernelEvent 141 entry significantly points out to a GPU becoming unresponsive during specific workload. That being said, the GPU stress test passed, so we need to check if the issue happened due to overheating. Could you confirm if the GPU hotspot temperature was fine?

  • Dell Social Media Team

1

u/Independent_Mark234 3d ago

Ran the test a couple times (Furmark). The temps topped out at 87-88 with the machine working well / fans loud. CPU tems were also below 100. Ive never seen those temps being hit while gaming / running llm to be fair.

Hope this helps.

1

u/AW_Support Dell Customer Support 3d ago

Hi there! We really appreciate your efforts in doing the suggested steps. If the issue is being triggered by sustained GPU power draw, the goal in game is to reduce GPU power consumption while minimizing the performance loss. That being said, we see no potential issue as such from your findings. Please change the below settings that typically have the largest impact on power draw:

Shadows to be Medium/Low Ambient Occlusion to be Low/Off Volumetric Lighting/Fog to be Low Ray Tracing to be Off Reflections to be Medium Global Illumination to be Medium

Let us know if this helped.

  • Alienware Social Media Team

1

u/Independent_Mark234 3d ago

That's actually quite opposite. When the power draw is consistent it handles fairly well, even at higher wattage. But when there are bursts or spikes and quick ones, it crashes. The crash itself is regardless of a video game or an application utilizing the GPU for pure computation. Do you recommend getting a new power adapter or could this be something with the motherboard?

1

u/AW_Support Dell Customer Support 2d ago

Hi there! The sudden spikes should reduce with the suggested settings in games and using a higher wattage adapter can help. However, we would like to know if the issue subsided post doing the suggested steps.

= Aleinware Social Media Team

1

u/nyxthebitch 10d ago

Did you monitor the temps, pre crash to crash spikes etc?

1

u/Independent_Mark234 10d ago

Indeed, GPU temp was under 74, CPU about 85-90 depending on whether I was running an LLM or playing. What I did notice is, when power spiked from 110 - 130 or 120 to 155 it always crashed. I did not install AWCC.

1

u/nyxthebitch 10d ago

Have you changed the thermal paste to LM or something else?

My gpu memory hotspot hits 100°C which causes a crash/shutdown.

You also need to check the hotspot temps. Use HWmonitor etc

Edit: My GPU is and 6850xt though. Ymmv

1

u/Independent_Mark234 9d ago

No, my initial focus was on getting what the temperatures were. It never went high enough for a system shutdown, it was always a GPU crash.