Pandas memory error

Question 1

I have a csv file with ~50,000 rows and 300 columns. Performing the following operation is causing a memory error in Pandas (python):

merged_df.stack(0).reset_index(1)

The data frame looks like:

GRID_WISE_MW1 Col0 Col1 Col2 .... Col300
7228260 1444 1819 2042
7228261 1444 1819 2042

I am using latest pandas (0.13.1) and the bug does not occur with dataframes with fewer rows (~2,000)

thanks!

Question 2

That wouldn't help here because I am using merged_df.stack(0).reset_index(1) in a pandas.merge operation....

Question 3

So it takes on my 64-bit linux (32GB) memory, a little less than 2GB.

In [5]: def f():
 df = DataFrame(np.random.randn(50000,300))
 df.stack().reset_index(1)
In [6]: %memit f()
maximum of 1: 1791.054688 MB per loop

Since you didn't specify. This won't work on 32-bit at all (as you can't usually allocate a 2GB contiguous block), but should work if you have reasonable swap / memory.

Question 4

Ahh, I am using Windows 7 64 bit, 8 GB RAM, but my pandas is 32 bit, could that be the issue?

Question 5

yep; you can install 64-bit python (and all packages), or use conda to do so. 32-bit has a 4GB addressable limit, but python requires contiguous memory, so that's too big to stack reliably. in my experience 32-bit has issues with anything > 1GB; 64-bit scales no problem however.

Question 6

@Jeff Thanks for the remark! I've been fighting for a good week with pandas to load only ~400MB of data in one dataFrame, when a list of smaller dataFrame instances, for the same total amount, can be loaded without a problem, and your explanation is surely the answer: I'm using Python in 32 bits, as our OSs at work are stuck on a Windows 32 bits. :-/

Question 7

As an alternative approach you can use the library "dask"
e.g:

# Dataframes implement the Pandas API
import dask.dataframe as dd`<br>
df = dd.read_csv('s3://.../2018-*-*.csv')

Jeff 130k21 gold badges223 silver badges189 bronze badges · Accepted Answer · 2014-04-21 23:30:03Z

5

So it takes on my 64-bit linux (32GB) memory, a little less than 2GB.

In [5]: def f():
 df = DataFrame(np.random.randn(50000,300))
 df.stack().reset_index(1)
In [6]: %memit f()
maximum of 1: 1791.054688 MB per loop

Since you didn't specify. This won't work on 32-bit at all (as you can't usually allocate a 2GB contiguous block), but should work if you have reasonable swap / memory.

Share

Improve this answer

answered Apr 21, 2014 at 23:30

Jeff's user avatar

Jeff

130k21 gold badges223 silver badges189 bronze badges

Sign up to request clarification or add additional context in comments.

3 Comments

user308827

user308827 Over a year ago

Ahh, I am using Windows 7 64 bit, 8 GB RAM, but my pandas is 32 bit, could that be the issue?

2014年04月21日T23:32:21.48Z+00:00

Jeff

Jeff Over a year ago

yep; you can install 64-bit python (and all packages), or use conda to do so. 32-bit has a 4GB addressable limit, but python requires contiguous memory, so that's too big to stack reliably. in my experience 32-bit has issues with anything > 1GB; 64-bit scales no problem however.

2014年04月21日T23:39:16.383Z+00:00

Joël

Joël Over a year ago

@Jeff Thanks for the remark! I've been fighting for a good week with pandas to load only ~400MB of data in one dataFrame, when a list of smaller dataFrame instances, for the same total amount, can be loaded without a problem, and your explanation is surely the answer: I'm using Python in 32 bits, as our OSs at work are stuck on a Windows 32 bits. :-/

2015年12月08日T15:09:53.467Z+00:00

CollectivesTM on Stack Overflow

Pandas memory error

2 Answers 2

3 Comments

Comments

Your Answer

Sign up or log in

Post as a guest

Post as a guest

Linked

Hot Network Questions

CollectivesTM on Stack Overflow

2 Answers 2

3 Comments

Comments

Your Answer

Sign up or log in

Post as a guest

Post as a guest

Linked

Related