Wow, this is fun. Numbers numbers and more numbers. Applying some simple standard diviation based algorithms had us at the score above for the middle 50% of the probe data. Unfortunately the upper and lower percentiles are proving nearly impossible to predict accurately without training our algorithms to the probe set only. So it seems we must do much better than a RMSE of 0.9685 for the stuff we can predict. Here are some numbers:
For the middle 50% we are at 0.9685, for the upper and lower percentile remainder we so far cannot get below 1.011, for an overall overage of 0.9968 which though a respectable RMSE it is not on the leader board though we may submit something soon just to see how our submission files work and maybe we just might get lucky? Wishful thinking I guess.
Anyhow, so far our Ubuntu Linux box (spider) is holding up well, our MySQL catalog for this project is now up near 15-20Gb, almost half of which are indexes to speed access to the data. I expect once we complete our cohort groupings we may need alot more disk space so I may need to swap that 40G for an 80G or larger, and maybe even getter more RAM, a faster CPU, and a faster hard drive while I'm digging around in there, heck I'll just get a near system to boot with one of them fancy SATA drives.
Our newest member D is trying more radical approaches, drilling into the data beyond the numbers as well working directly on the Linux box using Perl and other ~ix stuff. We downloaded the IMDB genre and other tables and may at some point do some joining and analysis using this data but we'll see.
Eat Fresh!!!
We are teamalrightythen, an amature team of computational guru's stepping up to the Netflix prize challenge.
Subscribe to:
Post Comments (Atom)

No comments:
Post a Comment