Showing posts with label Uninitialized memory. Show all posts
Showing posts with label Uninitialized memory. Show all posts

Wednesday, June 27, 2007

My 64-bit porting experiences - VI

6. Excessive code-reuse

We have always been instructed to reuse code, as it makes the code more maintainable. But it can sometimes be taken too far (too much of a good thing ?!).

6.1 Objects of different sizes in a union

Class A{
int getVal {return bVal;}
exp *getExpr {return bExpr;}
int hasExpr {return isExpr;}
union {
int bVal;
exp *bExpr;
}
int isExpr;
};

The object of this class ‘A’ can have a data-member that is either a constant value (bVal), or an expression (bExpr) if the value is not constant. When the value is not constant, ‘isExpr’ is set to ‘1’, and the user is expected to use the expression returned by ‘getExpr’ call.

In the following piece of code , ‘obj_a’ is an object of class ‘A’:

int tmp = 0;
exp *tmpExpr = NULL;
if (cond)
tmp = 1;
else
tmp = obj_a->getVal();

if(obj_a->hasExpr())
{
if(tmp == 1)
tmpExpr = createExpr(1);
else
tmpExpr = obj_a->getExpr();
}
else
tmpExpr = createExpr(tmp);

Incorrect values in a testcase, in 64-bit mode, were traced to this piece of code. Even though ‘obj_a’ contained expression (which represented a value other than ‘1’), and ‘cond’ was ‘0’, this code inferred the value represented by ‘obj_a’ to be ‘1’.

Reason: ‘obj_a’ represented an expression, so obj_getVal() should not have been used. But it was called, and it returned ‘1’. In 64-bit mode, the members of the union – int and pointer – have sizes 32 and 64 bits respectively. When the integer value was queried, the higher 32 bits were returned, which constituted ‘1’ in integer value [the pointer was 0x1_hhhh_hhhh]. This code worked without problems in 32-bit, because int and pointer have same size, and a pointer cannot be 0x1.


6.2 Overriding functions through implicit casting

This problem was encountered at many places in the code, in different forms.

A hash-table to hash pointers had been implemented, and widely used. The interface functions to insert new records, or find existing ones, accepted char** arguments (and dereferenced them to obtain the pointer). Later on, there was a requirement to hash integer values as well. Instead of implementing a new hash-table (or even writing new interface functions), the previous table and its interface was reused – by passing int* on the actual. When the interface functions dereferenced the pointer passed at the actual, the resulting value was int, and therefore not a valid address.

Tuesday, April 24, 2007

Memory, memory everywhere ...

... and not a block to link!!

A junior developer came to me in the morning, seeking help on calloc. I briefly described calloc and malloc to her, but I was somewhat doubtful of her requirements, so I asked what was the specific problem she was facing.

She had allocated a string using the malloc call, and was appending to the string using strcat. The result she got was junk characters.
str = (char *)malloc(N * sizeof(char));
for( i = 0; i <= m; i++)
strcat(str, arr_of_str[i]);


The answer lies in the behavior of malloc and strcat. The memory allocated my malloc is not initialized. So, when 'str' is allocated, it is filled with junk characters. The function strcat appends a new string to an existing string, and to identify the end of the existing string it searches for the null character ['\0']. In this case, strcat appended the new string [arr_of_str[i]] wherever it found a null character in the string 'str' - the initial characters remained as junk, and this is what she saw. In fact she was lucky to get away with a garbled string. Had there been no null character in 'str', strcat function would have written into the memory of some other variable [wherever it found a null character in the memory space adjoining that of 'str'], and caused a crash.

The fix was simply to initialize the newly allocated 'str':
str = (char *)malloc(N * sizeof(char));
strcpy(str,""); /* or alternatively, str[0] = `\0` */

Now, this sounds very obvious. But how often the obvious is overlooked, will perhaps be borne out by the fact that I had come across the very same problem not long back.

Tuesday, March 27, 2007

Beat me, whip me, make me use uninitialized pointers

Well, the title is just a catchy line I "borrowed" from a friend's custom message. The problem I am about to discuss does not have to do with pointers, but it indeed has to do with uninitialized variables.

The software product that I work on, is supported on three different UNIX platforms (Solaris, AIX, Linux), on different flavors of each of these. When a test cases starts failing on some of the platforms, especially on a random basis, it is fairly safe to assume that a memory corruption has happened. The primary software tool that we use to analyze memory corruptions is IBM Rational Purify.

A few days back some testcases in our test suite started failing due to missing messages from the log file - the failures were random, mostly on Solaris 9 and 10, and some times on AIX (almost never on Solaris 8 and Linux EE and OEE). I was almost certain that a memory corruption had been introduced in the code. What was surprising was that there was one particular message that went missing, and that the failures existed only in one stream, though it was not very different from two other streams, on which no such occurences were reported. But such is the nature of memory corruptions.

So, I ran Purify on one such testcase, but it reported no error.
Then, since I was fairly confident that it was nothing but a corruption, I tried Valgrind as well. Valgrind is a free software from GNU, available only on Linux (Purify is available for both Solaris and Linux), and it does not have a fancy GUI like Purify. But then, it does not have a fancy price tag either. [My primary development platform is Solaris, and the company buys Purify licenses, so my first preference is to use Purify, rather than any other tool.]
Valgrind did point out read of uninitialized memory - the value of a bit-field was tested to issue the message under analysis, and this bit-field was not initialized in some scenarios.

The interesting part to note here is why was the problem not reported by Purify, which is usually quite accurate - it owes to the way bit-fields are stored in a structure or a class object, and retrieved from the memory. When a structure (or an object) declares some bit-fields, these are packed together, and padded with empty bits to align the object at the word boundary. When the value of a bit-field is read, the OS reads the complete word, rather than the individual field. Purify works on the granularity of a word, so it will report an uninitialized memory read if some of the bits of the word are not initialized. Now, the empty bits that were padded for alignment will obviously ALWAYS be uninitialized; so to avoid false warnings, in the default mode Purify suppresses the uninitialized read messages in case of bit-fields.

For those who are familiar with Purify, the Purify error code for uninitialized read is UMR [Uninitialized Memory Read]. For bit-fields, the warning that is issued (and which is suppressed by dfault) is UMC [Uninitialized Memory Copy].