Skip to content
All library documents

Cleaning Malformed CSV Values Before the Augmented Dickey-Fuller Test

Article Quant Q&A · Author: TAN YONG SHENG

Summary

The document explains a data-format cause of a ValueError when running an Augmented Dickey-Fuller test with statsmodels. The reported workflow reads a CSV into pandas, selects a column, converts it to a NumPy array, and passes that array to the test. The failure occurs because at least one supposed numeric value contains an unexpected character, such as a question mark before a decimal number, so the data cannot be converted to floating point values.

The practical lesson is to inspect and clean the input series before applying statistical tests that require numeric observations. The example identifies a likely source in the CSV rather than an error in the adfuller call itself. It does not show a specific cleaning command, explain how to handle other malformed or missing values, or provide a validation procedure. The diagnosis is therefore useful as a first debugging step, but the appropriate treatment of invalid entries depends on the dataset and should be checked before testing.

Key ideas

  • Augmented Dickey-Fuller testing requires a numeric input series.
  • Unexpected characters in CSV values can prevent conversion to floating point numbers.
  • Inspect the source data when a statistical test raises a string-to-number conversion error.
  • The document identifies a likely cause but does not give a complete data-cleaning procedure.

Tags

Full text
# ValueError for adfuller test python


# ValueError for adfuller test python












[Solved] For the below code, I get a ValueError:could not convert string to float. I have tried to search from the forum but still not able to find the solution.

Hope for help. Thanks in advance.

[Solution] There is error for my data in csv file, e.g. there is some unknown format data like "?0.2" instead of "0.2" inside the csv datafield.

```
import pandas as pd
import numpy as np
from matplotlib import pyplot as plt
from statsmodels.tsa.stattools import adfuller

%matplotlib inline

df=pd.read_csv("./min-temp.csv", index_col=0, parse_dates=True)
df=df.iloc[:,0].values
dftest = adfuller(df, autolag="AIC")
```

The df result (https://i.sstatic.net/917nO.jpg)

Value Error (https://i.sstatic.net/yqB4Z.jpg) (https://i.sstatic.net/j51wK.jpg)

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.