Mars Pathfinder: What Really Happened?

You’ve probably heard of Mars Pathfinder, the NASA lander and its Sojourner rover that took up residence on Mars in 1997. At first, all was well, but after a few days, the lander suffered a series of total system resets. A watchdog timer detected a problem and forced reboots. [Vivek Bhageria’s] recent post looks into what happened and the eventual solution.

The root cause was priority inversion. This happens when a lower-priority task holds a resource needed by a higher-priority task and prevents the higher-priority task from executing. In this case, the operating system was VxWorks, but this can happen on any kind of priority scheduling system.

The simplified version is that there were two tasks that each used the same mutex to avoid stepping on each other. The highest priority was a data distribution task. The lowest priority task dealt with meteorology data processing. There was also a medium-priority task that handled bus maintenance.

The failure would occur when the low-priority task had the mutex and lost its time slot before releasing it. The high-priority task would then find it couldn’t take the mutex. However, the low-priority task didn’t get control again because of the medium-priority tasks waiting to run.

In effect, the high-priority task was now waiting on something that couldn’t become available until the lowest-priority task was able to run. That’s priority inversion.

Continue reading “Mars Pathfinder: What Really Happened?” →