Hi there, I am fairly new to programming and trying create a program to parse csv data into an array. I have been following an example I found on this site but I am running into a segmentation fault (even if I copy code directly). I am fairly sure the problem lies in the while loop, I just cant see it.


****Sample Csv Data****
1.0235, 2.1112, 1.9972
3.0139, 1.1876, 2.91785
1.9295, 2.1322, 2.4821

I realize its probably overkill for my purpose so if anyone could point me to a simpler solution that would be appreciated too.

#include <stdio.h>
#include <stdlib.h>
#include <string.h>
#define MAXFLDS 200     /* maximum possible number of fields */
#define MAXFLDSIZE 32   /* longest possible field + 1 = 31 byte field */

void parse( char *record, char *delim, char arr[][MAXFLDSIZE],int *fldcnt)
{
    char*p=strtok(record,delim);
    int fld=0;
    
    while(*p)
    {
        strcpy(arr[fld],p);
		fld++;
		p=strtok('\0',delim);
	}		
	*fldcnt=fld;
}

int main(int argc, char *argv[])
{
	char tmp[1024]={0x0};
	int fldcnt=0;
	char arr[MAXFLDS][MAXFLDSIZE]={0x0};
	int recordcnt=0;	
	FILE *in=fopen(argv[1],"r");         /* open file on command line */
	
	if(in==NULL)
	{
		perror("File open error");
		exit(EXIT_FAILURE);
	}
	while(fgets(tmp,sizeof(tmp),in)!=0) /* read a record */
	{
	    int i=0;
	    recordcnt++;
		printf("Record number: %d\n",recordcnt);
		parse(tmp,",",arr,&fldcnt);    /* whack record into fields */
		for(i=0;i<fldcnt;i++)
		{                              /* print each field */
			printf("\tField number: %3d==%s\n",i,arr[i]);
		}
	}
    fclose(in);
    return 0;	
}

Dani AI

Generated

A few practical fixes will make this robust and easier to extend:

  • Guard arguments and I/O: check argc before using argv[1], and handle long lines gracefully. If available, getline avoids fixed-size buffers.
  • When continuing tokenization, pass NULL to strtok/strtok_r, not '\0'.
  • Your sample has spaces after commas; trim leading/trailing whitespace around each field before using it.
  • Bound-check both the number of fields and field length to avoid overrunning arr.
  • If you actually want numbers, convert tokens with strtod and validate the result.

Example: parse one CSV line of doubles, ignoring spaces and handling bad fields safely.

#include <ctype.h>
#include <errno.h>
#include <string.h>
#include <stdlib.h>

size_t parse_doubles(char *line, double *out, size_t max)
{
    size_t n = 0;
    line[strcspn(line, "\r\n")] = 0;               /* strip newline(s) */
    char *save = NULL;
    for (char *tok = strtok_r(line, ",", &save);
         tok && n < max;
         tok = strtok_r(NULL, ",", &save)) {
        while (isspace((unsigned char)*tok)) tok++; /* ltrim */
        char *end = tok + strlen(tok);
        while (end > tok && isspace((unsigned char)end[-1])) *--end = 0; /* rtrim */
        errno = 0;
        char *ep = NULL;
        double v = strtod(tok, &ep);
        if (!errno && ep != tok) out[n++] = v;     /* accept valid number */
        /* else: handle/skip invalid field */
    }
    return n;
}

Note that strtok is destructive and cannot parse quoted CSV (commas or newlines inside quotes). For fully compliant CSV, implement a small state machine or use a CSV library following RFC 4180. See the C strtok reference for details and caveats at cppreference.

Recommended Answers

All 5 Replies

line 12: should be this: while(p) Your program gets segment violation when strtok() returns NULL.

>>I realize its probably overkill for my purpose so if anyone could point me to a simpler solution that would be appreciated too.
In C language that's about the best anyone can do.

I am not sure if this is needed or not, but most CSV files have a header record, which is essentially the names of the columns separated by commas. If you are parsing an actual CSV file, you will need to take that into account also.

> I am not sure if this is needed or not, but most CSV files have a header record, which is essentially the names of the columns separated by commas. If you are parsing an actual CSV file, you will need to take that into account also.

That may be true for some, but the vast majority of CSV files in use have no such thing. Unlike something like an XML file, CSV files have no way to tell you what they contain. The reading program must know what to expect when reading any CSV. Even the assumption that CSV is textual data, or fields separated by commas or tabs, and other stuff, is entirely that, assumption.

If you like I'll dig up an old CSV parser I wrote that can handle anything and port it to C for you. (Remember, it can read and write any CSV file, but, again, you must know what that CSV file contains for it to be useful!)

Ahhh thanks for that, a silly mistake.

Duoas if you have it lying around Id love to take a look

Thanks in advance.

Hey there, if you're still interested, don't give up.

I wrote it originally in Object Pascal, and I used a number of very powerful container classes that don't exist in C (...obviously). So I've been coding a replacement for them also... (It's kind of boring actually...)

I should be done pretty soon. Once I give it some basic testing I'll wrap the whole thing up for you and respond here again.

Be a part of the DaniWeb community

We're a friendly, industry-focused community of developers, IT pros, digital marketers, and technology enthusiasts meeting, networking, learning, and sharing knowledge.