Showing posts with label Purify. Show all posts
Showing posts with label Purify. Show all posts

Wednesday, October 3, 2007

Identity Crisis

Early this week, I got crashes in many testcases, on 64-bit mode, and only on Linux platforms (Solaris was fine in both 32- and 64-bit modes). Debugging such issues is a big problem, because they are almost certainly tricky memory corruptions, and debugging in 64-bit environment is seriously hampered by unavailability of proper support in debugging tools.

There are some specific techniques I follow to debug such corruptions or environment specific issues.

1) Reproduce the crash in an environment, where it is easier to debug.

Now, Workshop, the graphical debugging tool that I use on Solaris supports both 32- and 64- bit modes well. I like it for its user-friendly interface, and certain features like pop stack, that are not provided by other debuggers. But since I could not reproduce the crash on Solaris, I had to start my investigation on Linux. On linux, I use DDD, which is a GUI based on GDB.
I was in for a surprise, since the DDD did not load my executable. I checked the usual suspects - Build, Run arguments, paths, library paths etc. This did not help.
After asking around, a helpful team-member told me that there is a separate 64-bit version of GDB. I set my path and library path to pick this version, and was able to load the executable in DDD.

2) Compare the behavior on two platforms. Start with a top-level analysis, followed by a step-by-step comparison.

DDD showed me the point of the crash, but did not allow me to access the object, a call to whose method generated the crash. Morover, DDD does not honour the default values of function arguments, that one can provide in CPP code.

The line of code was something like:

if (obj && obj->hasProp1())
{
info = obj->getInfo(PROP1); // => the crash occured here
}

I ran the testcase on Solaris, but the execution did not reach the offending line!

3) Quick, high-level debugging through "printf" statements.
I put a "printf" statement before the "if" statement, to display the name of the object 'obj' and the value returned by hasProp1() call. I compared the output on Solaris and Linux, and the output was exactly same! This value was '0' on both platforms, and how did the code entered the "if" condition, is a mystery I could not solve till the end.

4) Use memory analysis tools. Try different tools, as it frequently happens that one tool catches the problem that the other one coud not.

I ran Valgrind on Linux, but it did not show any errors. Valgrind is available only on Linux, but it has no build-time requirements.

One point to note is that one can catch corruptions equally well on any platform, since the problem exists everywhere, even if it does not manifest in certain scenarios.
Purify was initially available on Solaris only, though now it is available on Linux as well. So I'm more comfortable using Purify on Solaris, and also in the past, I've had problems using Purify on Linux. So, I tried to run Purify on Solaris. However, I was not able to run the 64-bit Purify build - it crashed even before executing the first line. I tried build/run several times a number of times, with the same outcome. After struggling with it for a long while, a thought struck me - I was running it on S10. Our buid platform is S8, but with the same build, we support run on S9 and S10 as well. I had been running non-purify build on S10, and I continued to do so with the purify build. So, in yet another attempt, I ran the testcase on S8. This time it did run! 64-bit purify build does not run on S10!
However, the run still did not help me much - it kept delivering SIGBUS infinitely. I loaded the purify build in Workshop, and put a breakpoint in the function purify_stop_here - this function is a debugging hook provided by Purify - it is called before Purify reports a memory violation. I got the point where the SIGBUS was reported, and analyzed that part of the code, but it only led me a wild goose chase. But well, that is life!

5) When all else fails, ask around!

I asked many people in the group if they have debugged 64-bit binaries on Linux. Finally, one person told me he had done so, but he does not have much faith in DDD. So he used GDB on the shell. He also gave me a pointer to the version of GDB he had (successfully) used. This was another version of 64-bit GDB!

6) Revisit steps

Meanwhile, I had also made a Purify build on Linux. I ran this Purify build with the new version of GDB. Purify reported an ABR (Array Bound Error) on the same line that DDD had reported earlier. I put a breakpoint in the function purify_stop_here, and continued till I reached the point of ABR error. This time, the version of GDB correctly reported the name of the object. I moved to Solaris Workshop once again, and looked for this particular object. Having the name of the object made it quite easy this time. And what I discovered was absolutely stunning - the object was actually of the base class, while the properties we were querying on it were functions and members of the derived class. The object had been return by call to a function, which could return either kind of object. The fix was, of course, extremely simple - I modified the "if" condition to:
if (obj && obj->isDrvCls && obj->hasProp1())
{
...
}

The class declarations were as follows:

class baseCls
{
private:
...
char *name;
public:
...
char *getName() { return name; }
int isDrvCls() { return TRUE; }

}

class drvCls : baseCls
{

public:
...
int hasProp1() { return type_ & PROP1_MASK; }
listCls *info() { return lst1_; }
infoCls *getInfo(int propType) { return info()->findInfo(propType; }
int isDrvCls() { return TRUE; }
private:
...
int type_;
listCls *lst1_;
}

Tuesday, March 27, 2007

Beat me, whip me, make me use uninitialized pointers

Well, the title is just a catchy line I "borrowed" from a friend's custom message. The problem I am about to discuss does not have to do with pointers, but it indeed has to do with uninitialized variables.

The software product that I work on, is supported on three different UNIX platforms (Solaris, AIX, Linux), on different flavors of each of these. When a test cases starts failing on some of the platforms, especially on a random basis, it is fairly safe to assume that a memory corruption has happened. The primary software tool that we use to analyze memory corruptions is IBM Rational Purify.

A few days back some testcases in our test suite started failing due to missing messages from the log file - the failures were random, mostly on Solaris 9 and 10, and some times on AIX (almost never on Solaris 8 and Linux EE and OEE). I was almost certain that a memory corruption had been introduced in the code. What was surprising was that there was one particular message that went missing, and that the failures existed only in one stream, though it was not very different from two other streams, on which no such occurences were reported. But such is the nature of memory corruptions.

So, I ran Purify on one such testcase, but it reported no error.
Then, since I was fairly confident that it was nothing but a corruption, I tried Valgrind as well. Valgrind is a free software from GNU, available only on Linux (Purify is available for both Solaris and Linux), and it does not have a fancy GUI like Purify. But then, it does not have a fancy price tag either. [My primary development platform is Solaris, and the company buys Purify licenses, so my first preference is to use Purify, rather than any other tool.]
Valgrind did point out read of uninitialized memory - the value of a bit-field was tested to issue the message under analysis, and this bit-field was not initialized in some scenarios.

The interesting part to note here is why was the problem not reported by Purify, which is usually quite accurate - it owes to the way bit-fields are stored in a structure or a class object, and retrieved from the memory. When a structure (or an object) declares some bit-fields, these are packed together, and padded with empty bits to align the object at the word boundary. When the value of a bit-field is read, the OS reads the complete word, rather than the individual field. Purify works on the granularity of a word, so it will report an uninitialized memory read if some of the bits of the word are not initialized. Now, the empty bits that were padded for alignment will obviously ALWAYS be uninitialized; so to avoid false warnings, in the default mode Purify suppresses the uninitialized read messages in case of bit-fields.

For those who are familiar with Purify, the Purify error code for uninitialized read is UMR [Uninitialized Memory Read]. For bit-fields, the warning that is issued (and which is suppressed by dfault) is UMC [Uninitialized Memory Copy].