Showing posts with label Access. Show all posts
Showing posts with label Access. Show all posts

Friday, May 20, 2016

Richard Strauss, and Data.

This past week, I've managed to get in to work fairly early almost every day so far (Monday- Thursday); I think Wednesday was the only day I came in later, and that was intentional. I've also been getting up earlier, and the time before work has been used to work on the database. I've made some good progress on the few boxed sets of audio CDs that we own, which is where Richard Strauss comes into the picture.

Before I go any further, I should give you a bit of my musical backstory, kind reader. You may have surmised from previous posts (especially the most recent one, Axl Rose, and Data,) that I'm more Rocker than Opera-Goer. This is 100% correct. However, my musical tastes are fairly diverse... Yo-Yo Ma, The Beatles, Chevelle, C.W. McCall, Air Supply, Deep Purple, Newsboys, Sibelius, 80's hair bands, choral, etc. I firmly believe that any music that is good should be played loud when possible.  I like some blues, classic Motown, and a smattering of jazz. I'm okay with the various forms of trance, and dance music- if it's something that catches my ear. I can even deal with disco these days. Things I pretty much have no interest in are rap, hip hop, opera and death metal.

So, what's with Strauss?

Well, the album pictured above is a boxed set I picked up at a library sale for a solitary U.S. dollar. Cheap-value-SCORE! It's a three disc set, with a booklet nearly as thick as the CD case. The 330 page "booklet" has all sorts of details, not merely about this opera, but about this particular recording, its cast and conductor. It also has the complete text and lyrics of the opera in French, English and German. When I purchased it, I had decided that the weight of the boxed set alone made it a good buy... little did I know that this was a "reference" recording, one by which all others are to be judged. Its also conducted by Herbert von Karajan, a legendary conductor.

Still, what's with Strauss?

Data entry, pure and simple. As this past week has seen a renaissance of Forty-Two, I decided to tackled boxed CD sets. The problem is that this particular set has sixty-two tracks, all of which have German titles. I'm slow enough at data entry without having to import special characters, so I did what any reasonable human being would do: I looked up the recording on Amazon, copied the track list and pasted it into Excel. From there, I copied and pasted each track into Access. That's where I stopped with music- I still have three boxed sets to go, and then it's on to albums.

In other data news, I've got what appears to be a workable solution for my internal Lego part number. It's fairly lengthy at seventeen characters, and from all appearances, this should be sufficient. I've begun the data entry on this, and tried a few trial sorts. So far, everything looks good, and this is officially stage 2 of the Peeron normalization. I still need to add dimensions and clean up the text descriptions before importing it into a table.

As always, I am hochspeyer, blogging data analysis and management so you don't have to.

Sunday, May 15, 2016

Axl Rose, and Data.

It's been busy at work lately. And the rain has been frequent. And... it's Springtime in my little corner of the world. Consequently, our lawn got a "little" out of control.

(Spoiler alert: this post contains references to Guns 'N' Roses and their songs... 
and a few other musical references!)

Jennifer was kind enough to help out and mow the front lawn, which is what everyone passing our home sees. The back... well, that's another story- and my bailiwick.

My time had arrived. It had not rained for over a day. I looked out the back window and saw the grass blowing like Dust in the Wind. As my gaze landed upon the compost bins, I had to rub my eyes because I thought I saw a black tophat sitting on top of one of the compost bins, and a bandana on top of the other. I shook my head, blinked and looked again. Without warning, the yard had gone from broad daylight to a starless night. The yard's verdancy had turned to a monochrome with rough-cut video quality- I kid you not!. Our neighbor's white fence had been replaced by a gaggle of Marshall stacks, spotlights were illuminating the compost, and Slash, with trademark shades, ciggie and Les Paul, and partially obscured by the output of several smoke machines was furiously spewing out riff after riff, while Axl crooned, "You're in the jungle, baby" as he seemed to float over the grass. I covered my eyes, and shook my head. When I looked again, the moment was over. My yard was green, and back.

Seven Marshall stacks- fourteen cabinets and seven heads 
 Sheesh! Less Monster, more cowbell... maybe.

So, I cut the grass. Yes, it was long- long overdue for cutting. I'm not certain how long it took to complete the task, but I had to make several trips to the compost to empty the mower's bag. When that was done, I decided to tackle another yard task: the ivy. And, to be fair to the myriad of horticulturalists, botanists and assorted green thumbs in my audience, I'm not really certain what this plant is. It has dark green leaves, a woody stalk, and will sprout roots along the stalk. It's also quite capable of climbing. So, we've decide that it must go.

I got most of it off of the red mulberry. Now, technically the mulberry should be a bush, as it has multiple trunks. This one, though, is ~30' (over 9 meters) tall, so I call it a tree. Many websites also call it a tree.

Our Mulberry last Winter

I put the mower back in the garage, and grabbed a pair of leather gloves, a few yard waste bags and some clippers, and proceeded to tear into the ivy. For reference, the bags are constructed of a double walled heavy brown paper, and are approximately 16" x 12" x 35" (40.6cm x 30.5cm x 88.9cm). I filled five of them, and I'm probably only about one-third of the way done. This was sometime last week, and since then, we've had rain at least four times, so the next time I tackle this project, the stubble from the first round should be visible, which should make the next phase a bit easier.

Finally- data! (Sorry, no data pictures!)

On Saturday morning (May 14th), I FINALLY completed the 1st phase of the Lego Peeron data normalization! Speeds and feeds are appropriate here, so here we go!

The original Peeron database I'm using is from March 19, 2012. It contains 18,510 parts (rows). After the first passes through the data- in which the data started out as a text file, and then was converted to an .xlxs file, and then the data had its initial cleansing where a "base" part number was created- the data ended up totaling 16,218 parts. It should be noted that this not a fixed number, as there are more parts that need to go to the "Stickers etc" worksheet. I also want to emphasize that I am cleaning data and not normalizing; after all, all of this data still resides in an Excel file, so it is still a flat file. My next task is to impose order onto this file, so that will require the creation of a unique part number. This part number will be the primary key once it is imported into Access.  

Before any of this happens, I need to come up with a standardized format for my unique part number- currently called the "DB_Tracking_Number". Ugly, but good enough for now. I can't really automate this, once again because of the lack of standardization in the Peeron data, and my attempt to impose my own personal spin on Lego organization.

So, for now, I have ~ 16,000+ part numbers to create.

As always, I am hochspeyer, blogging data analysis and management so you don't have to.



Saturday, March 5, 2016

Working with someone else's data

At a certain point in one's data career (plain English: pretty near every day, and with nearly every piece of data one normally touches!), you will have the "opportunity" to work with data someone else has prepared, or at least accumulated or aggregated or in some way modified. As I create Forty-Two, "most" of the data I will be using is data which I have entered (created) into a table. There is one notable exception to this, though, and it is for me a moderately to severely painful one: Lego parts.

As mentioned in a previous post, my experience tells me that the best place for my Lego data is in a spreadsheet- at least initially. And although calculations can be done in Access, Excel is much easier for me. And, they can be linked at a later date.

However, there is a HUGE caveat to the Lego data. The title refers to someone else's data. In my case, I'm using peeron.com's parts list. It's the best one that I've found, and many other AFOL (Adult Fan of Lego) sites use it.as a parts reference. The problem with the list and site is that they're horribly out of date. The list has a very standard naming convention, and the last update was nearly four years ago at the time of this writing. However, as it is the best, easiest to use and most complete list currently available, I've decided to use it. I can update newer parts as I find them on other websites.

The other issue with the Peeron data is that it is a .txt file. Not bad when importing to Excel- just copy and paste. However, to make it usable, I have to manually edit the 18K+ rows of data. As I'm in no great hurry to finish this phase of the project, it's not too much of an issue. Still, .... I know a guy (as they say) who might be able to help. More on that later.

The final issue with Peeron is that its creator strove to make it a very complete parts list. As such, there's a great deal of data which I actually don't need- stickers and superseded part numbers are two types which come to mind immediately. So, if my guy can fix my primary issue, the other issues will be much easier to deal with. If not, then I still have a great deal of work ahead of me. UPDATE: I'm not really surprised- but also not horribly disappointed- that we were not able to fix the data.

Before I forget, I wanted to post a brief update on the blog itself. I'm not quite sure when I did this last, but I think it was around the time the blog hit 10K viewers. As of today, the blog has over 15K viewers in fifty-five countries on six continents (c'mon, Antarctica!). Africa is represented by three countries, Asia by seventeeen, Australia by one, Europe by twenty-seven (I'm counting Russia in the Europe column rather than Asia), North America by six and South America by one. What's most amazing to me is that of these fifty-five countries, I could only count seven where English was either the official language, a dominant language, or one of a group of commonly accepted languages.

To each and every reader- THANK YOU!

Last: a small compilation of my blogs dealing with data (for those who are interested in how a small-time operator handles data)

http://hochspeyer.blogspot.com/2016/02/you-said-this-was-about-data-analysis.html
http://hochspeyer.blogspot.com/2015/06/data-science-pt-1.html
http://hochspeyer.blogspot.com/2016/02/a-database-against-rules.html
http://hochspeyer.blogspot.com/2015/05/data-defined.html
http://hochspeyer.blogspot.com/2015/02/forty-two-v7-or-so.html

As always, I am hochspeyer, blogging data analysis and management so you don't have to.



Wednesday, February 3, 2016

You said this was about data analysis!

Regular readers are familiar with the tagline that ends every one of my posts, "As always, I am hochspeyer, blogging data analysis and management so you don't have to." There's a bit of truth there, as well as a bit of attempted humor.

For starters, I've been using "hochspeyer" as my online persona for years. "Hochspeyer" (upper cased H) is a village in Germany where Jennifer and I lived for a few years.  My online name (hochspeyer, lower cased h) is something I've been using for years. And although the topic of data doesn't come up as frequently as I'd like, I do try to feature it when possible. One problem I have is that I don't do as much analysis as, say, when I was actually employed as an analyst... for 15 USD/hr.

Pregnant pause for effect.

Yes, I once made only $15/hr working as a data analyst (this might be lots of money in some parts of the world, but in the greater Chicago area, its not- especially when you're feeding, clothing and providing housing for five persons). According to payscale.com, it should have been (at a minimum) 39K per year. According to an inflation calculator, the USD inflation rate between then and now is ~17.6%, so I was roughly 2K under the base salary for an analyst, or roughly 5%. However, at that point in my life, it was the only gig I could score, and the plus side was I learned quite a bit about basic data analysis, Excel, Access, customers, internal and external deadlines, and adding value to both your output and your position.

As a temp, I did analysis for $12/hr. My boss there was a lady who was the Quality Manager. The plant was a contract manufacturer of aluminum die castings for the automotive industry. Although they produced products for all three of the "Big 3" (General Motors, Ford, and Chrysler), their main customer was General Motors (GM). This was when Saturn was an exciting GM product line. The company had made a significant investment in equipment that was specifically designed for small cars, such as the Saturn line. Unfortunately, the Saturn line had disappointing sales for GM, and was ultimately scrapped. I apparently was doing a good job there, however: there were two rounds of layoffs in the company before I was let go. My value add there wasn't special- I integrated Excel data into an existing Access database. I also had the temerity to give reasons when they asked why I couldn't extrapolate results (when one has neither data or experience, one is up the creek without a paddle in terms of validity!)

With overtime, computer crashes and life in general, I haven't had much time to think about data as of late, much less do any analysis. However, the further I get from 10K blog viewers, the stronger the pull is to do a bit of analysis on blog hits to date.

So, here we go: blog speeds and feeds.

I started writing this blog in December of 2012. As of January 2016, I've published 196 posts to this blog- roughly five posts per month. As I'm no professional writer, I've got to say that this paltry volume is a challenge to maintain.

I've been tracking numbers from Day One, and they break down like this:

As far as continents go, only Antarctica has zero readers- Antarctica is 7th.

South America comes in 6th. With the limited amount of data I get back from Google (it's free, so I'm not complaining), I can only guess that language may be a factor in the readership here.

In terms of continents, Australia has the fewest number of countries... 1, and it is predictably 5th in the ranking. However, I think a few more countries could probably be considered a part of Australia for statistical purposes- New Zealand comes to mind immediately, but that's about it. If anyone has any ideas, please comment. The rest of the continents are as follows:

Africa (4th) is actually tied with Australia in terms of total readers, but these readers are distributed among a few countries, so I have to award the spot to Africa.

Europe comes in 3rd, which is slightly surprising because of the widespread use of English there. Twenty-seven countries represent Europe.

Asia is represented by seventeen countries and comes in 2nd, mainly because of Chinese readers.

My top market, again not a surprise, is North America, with a bit over 75% of my global total in readership. 

That's it from here for now. Before signing off, I'd like to extend special e Day (2/7) greetings to Prof. Diego Kuonen!

As always, I am hochspeyer, blogging data analysis and management so you don't have to. 



 




Saturday, April 25, 2015

Data, evolving

Quick- fill in the blank: I do some of my best work in ______________________.

If you filled that blank with "the bathroom" congratulations! You, Freddie Mercury and I think alike. I believe Freddie Mercury claimed to have come up with the idea and melody for "Crazy Little Thing Called Love" in ten minutes while soaking in the bathtub in Műnchen (Munich), Germany, and the same thing happened to me Thursday evening!

Well, sort of... I didn't actually come up with a #1 U.S. Billboard chart single for a supergroup in a hotel in the Black Forest in Germany, but I think I've come up with the way I want to get my Lego data into Access.

A few years back, I had designed a flat database for my Lego collection in Microsoft Excel 2003. It was pretty nice and fit my needs at the time quite nicely. It was composed of several pages where the data was entered, and then a couple of pages which gave some bird's eye level analysis via links to the data pages. The project was scrapped mainly due to my impatience with data entry and my inability to actually get an accurate count of all of those elements- besides, they are much more fun to build with than count.

So, my plan is to go back to that format, but instead of merely having links within the workbook, have links from the Excel 2007 workbook to the Access 2007 database. And why 2007 rather than 2010 or 2013? Expediency: the Legos are close to a machine in the Secret Underground Lair (SUL) which runs 2007. The SUL redo is going very slowly, but if I am successful in liberating space for my laptop in my corner of the SUL, then it will be done on the 2013 versions of Excel and Access.

I'd like to get started on this first- creating the spreadsheet with the data, and then importing the data into an Access table and then (hopefully!) linking each data cell in Excel to its corresponding Access field so that the database can be updated via Excel. Its been some time since I did this, and it was on the "pre-ribbon" interface, so it looks like I'm going to be relearning some Excel tricks.

In other data news, Jennifer and I have been walking more as the weather has become a bit more pleasant. My replacement pedometer (identical to my expired Omron model) is chugging along, but its going to be at least another month before I have daily month-to-month data available.

As to my new glasses: the lady who fitted me for frames was slightly surprised as to how quickly I adjusted to my new spectacles, although I think I'm still breaking them in. She also said something that caught me off-guard: I wear my lenses lower than most folks.

 I've never worn gradient lenses before- I've worn reading glasses since the age of eighteen and I'm now 50+; my new specs have three prescriptions per lens: top for distance, middle for computer work and bottom for reading. Although the glasses really do a good job (especially while driving), they are singularly ineffective in the place where I need them the most: computer. Here's the deal: before I went in for the exam, I had made some careful observations and mental notes about my work environment and how I typically used glasses, as well as my concerns about night driving. The gradient lenses the optometrist came up with (pretty much six prescriptions- three per lens!) are fantastic for general purpose applications (life!). My hope in getting this type of lens would be avoiding a second pair of glasses. To quote "Mad Dog" Tannen, "You thought wrong, dude." (*BLAM*) I was very quickly able to pick up the reading area of the lenses, and the distance area (which I don't really need but seems to be beneficial in low light/nighttime conditions) also "snapped" into place pretty quickly.  I really need the computer part, though, and this is where I'm experiencing issues. Simply put: the lens cannot focus on the whole screen- even when I push the frame up tight to my face for the best focus. In the "good ole" 15" CRT days, these glasses would have been miraculous; with the 23" CRT's I have at work and at home, the lenses are just not up to it. And honestly, at home it really doesn't matter, because I'm generally not looking at the whole screen. At work, though- I do a lot of page layout and composition, and I generally NEED to take in the whole 23" diagonal screen at a glance- and more often than not, at a 90 degree offset. So, after only two days, I was back getting fitted for single vision computer glasses! (Just a hint here: when your livelihood depends upon a durable medical appliance- even if your employer or insurance doesn't pay for it: BUY IT!) The good news is that the second par of specs was massively discounted, and I should have them in time for my return from vacation.   

One final note before I call it a night: my twitter account @CjoelHarrison grew 10% in twenty-four hours!

As always, I an hochspeyer, blogging data analysis and management so you don't have to.


Tuesday, February 10, 2015

Forty-Two, v7 (or so)

I've been spending a fair amount of time on Twitter (@CjoelHarrison) lately... so much, in fact, that other elements of my e-life have been suffering. Case in point: this blog. In January of 2015, while I was establishing my presence in the twittersphere, I did not post a single blog. Possibly even more tragic, my playing time on Rail Nation has suffered tremendously. And don't even get me started on my lack of quality Steam time... I'm not certain when the last time I played Sid Meier's Civilization V was.

In any event, here I am once more heralding the latest iteration of my pet database project, Forty-Two. This may indeed be Version Seven; then again, I've not kept track, so it could be any number greater than three. This iteration, though, is different from previous ones: it was not necessitated by a hardware crash or a software glitch. Rather, it is simply a fresh start with a few ideas to make a (relatively) large dataset more organized.

So, I suppose the logical question would be this: why would a home user with no database training want to even try to build an Access database in the first place?

I suppose I could just write it off as an Aspie thing- a project that goes on forever without a realistic chance of being completed (I'm not sure that's exactly an Aspie trait!) It could be that I haven't learned MySQL yet, or haven't even tried NoSQL, and let's not even bring Hadoop into the discussion!

The reality is a lot more pedestrian. And practical: I want to know what I don't know.

In plainer speech, I want to learn how to build a small relational database from the ground up. I want this database to be useful to my family. And, possibly the most important reason for building the database: should anything bad happen, I'd like to have it as a record for an insurance claim.

So there, that's what drives me to keep on trying to build Forty-Two, the database that answers the question of Life, the Universe, and Everything.

The other side projects that have been taking up my time are coding. Professionally, I'm a programmer. The interesting thing about that is that even if I spent the majority of my time programming, the world would probably not recognize me as a programmer, because the tools that I use are very specialized and are more related to composition and layout than programming- and when we DO actually write some code, there is absolutely no way it con be confused with a Turing complete language. Still, some of these tools DO have actual programming, and to help myself to better understand these areas, I've taken to studying the Python language.

Gasp- it's official now: yours truly is studying Python.

print("That's all for now folks!")