Showing posts with label Lego. Show all posts
Showing posts with label Lego. Show all posts

Monday, November 19, 2018

The Lego database, reboot

The "reboot" seems to be something of a thing in Hollywood these days, and so it only follows that Life would imitate "art" as I reboot my Lego database.

For those who have been following this blog for some time, I think you may be aware of my Lego database project. It is one component of "42", my personal, ultimate repository of those possessions of mine which I have chosen to catalog.

I've just started using Microsoft Office 365, and the Lego data came from rebrickable.com. Their dataset is very comprehensive, but I don't think I've ever downloaded ANY Lego dataset that was usable by me as downloaded... and this data is no exception.

I mentioned that the data is quite comprehensive, which means it contains things which I don't necessarily need or want. For example, I do not care about any decals, paper or cardboard items, books, or even certain parts of the Lego product line. So, these need to be removed. The data also appears to come from more than one source, and formatting is necessary. Punctuation needs to be removed from many entries, and part names need to be standardized. Some data needs spelling changes- there are cases of the Queen's English being used, so "windscreen" and "tyre" need to be changed to "windshield" and "tire". These are the major changes that need to happen before I can even think about exporting to an Access database. And these are just a few examples.

But, I digress. Here are the numbers as of December 4th.

When I first started out, there were approximately 29,000 records. As of last night, after culling out items which I knew I would not be inventorying, I had 27,164 unique items. When I deduped the group to include ONLY the base part numbers- excluding all decoration variants, I ended up with a working list of 8,747 unique base part numbers. As I copy the part descriptions to the new inventory list, they are being further culled.

I am now at the point where everything must be done by hand. I find myself going back and forth between rebrickable and my flat database, verifying that the part number in the list is a part that I (may) actually own... and want to count. At a certain point, some of the data is subjective, and even though it is valid, I will count a complete assembly rather than, say, a special tile, wheels and tires as separate parts.

As always, I am hochspeyer, blogging data analysis and management so you don't have to.


Sunday, June 3, 2018

Music Without an Agenda

I can't speak for anyone else, but when I have music going  on my computer, its generally for a purpose. I NEVER shuffle my music, because my music collection- although conventional- is rather eclectic. Now, before you walk away, realize that this admission is akin to admitting that "jeans day" at work is having the ability to wear jeans. Just like everyone else.

However, THIS blog is about being contrarian in the face of conventionality or trends.

After a LOT of years, I've decided to build a computer for myself. In the past, we've built computers for Jennifer and Mr. T, and now its my turn.  I suppose the first question is, "Why build a computer"?

Well, for starters, because I CAN.  On a certain level, PC building is a lot of fun. On a more personal level, its a bonding experience for Mr. T and I.

At this point, I need to make a note about this particular build. After having three PCs die on me during this year, I decided it was time for a NEW PC. So, I spent a month or two researching cases.

Eh?

Yes, you read that correctly. I spent a few months researching cases.  Why, you may ask.

Well Jennifer's PC was one of our first builds, and it was a "bare-bones" kit which included everything needed to have a running PC- all that was needed on the purchaser's end was a few additional parts, an operating system and assembly. As Mr. T puts it, "building a computer is a lot like making something out of Legos. Very expensive Legos". Then, he specced out an I5-based computer for himself, and we built it. In the interim, I've had several PCs die on me (apologies to readers who less blessed- this is TRULY a 1st world "problem"). Having said that, though, by "1st world standards" my family is comparatively not particularly wealthy- we just look for the best value we can find.

Of the 3 PCs that died on me this year, two were ancient boxes which I purchased some years ago as refurbs- and they were at least a few years old then. And honestly, they served me well. The third, however, was my laptop, which suffered a harddrive (HDD) crash. After that, the power supply died. This was the only one of the three worth salvaging. As much as I really like my Lenovo laptop, I need a desktop- something with a large screen (or two) which I can crunch numbers on and look at data. And play some games.

So, after a few months of research, I had narrowed my choice of cases down to a few. Mr. T had envisioned this build as a small case with a small ITX motherboard. After doing some research, I decided that the ITX had too many limitations, and  decided to "go big or go home".

In the course of my research, I eventually narrowed my choices down to two, both Thermaltake offerings: the Core X5 tempered glass, and the core X9.  Both are MASSIVE cases by any standard.

In the end, I decided to go with the  Thermaltake Core X9 Core Snow edition.  Now, the picture at the right shows the case in its shipping carton. For reference, we have four cats, and the cat tree behind the case is six feet (~2M) tall.

Snow Edition? Yes, I paid extra to get a white case... ~30 USD, actually. And why? Jennifer and I had this discussion, and after all of his resistance, Mr. T stepped up in my defense., saying , "This is his build. He's going to be the one looking at it for the next several years." And beside that, there is a very practical reason for white: when building, its harder to lose things!

So, I bought a white case. But, unlike a lot of the DIY PC folks who put their builds up on youtube, my pockets are neither deep or wide. The case is just the 1st step. One of the things that Mr. T gave me grief over during the initial planning stages was the inclusion of a Raspberry Pi in the case. I didn't have a specific scenario, but thought that the inclusion of a Pi in a large case would be schweet!





Saturday, October 15, 2016

Forecast for tomorrow: sore everywhere, with scattered achy

There was one thing that I had hoped to get done on my vacation, and that was yard work. We're officially in early autumn here in greater Chicagoland, but the casual observer or visitor would not have surmised that based on the temperatures, it was late August. According to some websites, the area averages nine days per month with high temperatures over 80 degrees F (26 C), but this year we had fourteen. Only the last three days of the month saw highs below 70 F (21 C). We also had thirteen rain or thunderstorm days. So, not a lot of yardwork got done in September.

Earlier in the week, I mowed the front yard, but it was hot so I left the back for the following day. The next day, it rained. Of course. I swear I cannot make this stuff up.

The following day was warm and breezy. I did the back. Yay... the first part of my work was done. The next part, though....

Today, Saturday, started off as the perfect autumn day for yardwork. While Jennifer prepared a much anticipated breakfast of pancakes, sausage gravy and effs over easy, I made a cup of tea and got out of the way. Schwarz came out and sat down on the picnic table. I savored the moment, taking in the cool humid air while catching up on Words with Friends and sipping my tea. The local high school (American) football team was at home, and occasionally I'd hear the spectators cheer or the band strike up a tune. As I'd completely caught up my games and the tea was gone, I went back in and had the most incredible breakfast!

The next hour or so was taken up by a long overdue project: removing a line of concrete blocks I had installed many years ago to separate a narrow strip of garden from the lawn. Once they were removed, I backfilled the resulting trench with compost. Which means that... most of the yardwork is done!




In Lego/data news, the Ferris wheel is moving along slowly, as I had anticipated that it would. I experimented last night with creating supports for the spokes, and I have a promising system... pics to follow if it works!



In other Lego/data news, Sensei Wu has an upgraded Bō. Astute observers are probably wondering at this point why I'm messing with Minifigs when they are clearly not an area of interest for me?

Well, for starters, Wu is a cool-looking character, and his hat is a fairly rare part (yes, I'm one of those AFOLs [Adult Fans of Lego] that will buy a set based on parts rather than the model). Wu seemed to beg for an upgrade, though, and the best one I could think of was to turn the Bō into a scythe. No problem- Mr. T and I eventually came up with the True Neon Green x346 part, but the process of finding the part was tedious. I probably have somewhere in the neighborhood of 50-100 of these in various colors, but none of them are in a dedicated parts drawer.  We located the green one in a drawer dedicated primarily to Bionicles, but with so few Bionicles currently assembled, the drawer needs to be sifted through. An inventory and better parts organization would have probably saved me ten minutes.

As always, I am hochspeyer, blogging data analysis and management so you don't have to.

Friday, October 14, 2016

The Lego MOC, and inventories

For those not involved with Lego, I present a cautionary tale of despair and hope. I've taken this week as vacation (Mon, Oct 10 through Fri, Oct 14). I had a few ideas for some things I might like to do, but for very good reasons most of them did not happen.

The one thing I did get working on, though, was my Lego Ferris Wheel (its a MOC, pronounced "mock"; MOC is the acronym for "My Own Creation") .

Let me say that a Ferris Wheel is not a Lego project to be undertaken lightly. For starters, Lego and circles are not natural pairs. Still, as in real life, a Ferris wheel can (and is) made out of things which are not naturally round.

For the record, this is my second Lego Ferris Wheel. I can't claim with any certainty that my first one predated any "official" Lego sets, but that first one was a very large, motorized and chain-driven model. Fairly conventional, actually.

I want my new wheel to be a bit more visionary, cutting edge- but to still be a traditional Ferris Wheel, adding a bit of fun that might be doable in real life: contrarotating sets of wheels. I'm still working this out, and believe or not, the biggest challenge out of the gate is designing the hub for the wheels.

These photos show a very rough mockup of what my original idea was. My biggest concern when I first conceived the idea was that I would not have enough parts for the hub(s). And herein lies the data part of my problem.

I believe I have often mentioned that I was working on a Lego inventory. While this is true, a project of this scope requires a somewhat more complete inventory than I have ever had. Fortunately for me, in lieu of a complete inventory, I have some pretty nicely organized parts drawers. Parts drawers?

Yup. Most of my Lego collection is currently housed in one of two types of plastic cabinets. There are nine small ones, which are highly organized (generally one or two parts per drawer), and also nine large ones. The large ones tend to be much less organized, with the exception of the drawers which house basic pieces, which end up in quart or gallon-sized Ziploc© freezer bags.

That's all for now. Hopefully next time I will have more Ferris Wheel progress to report.

As always, I am hochspeyer, blogging data analysis and management so you don't have to.

    

Sunday, July 24, 2016

Professional email addresses... and other terminally flawed data

I'm sorry. That's a pretty bloody lame title. And I'll also forgive you if you didn't read my previous post on this subject... which had such a clever title that I can't find it either!

Every so often, though, I read something that makes me wonder: why did the writer go into business at all?

Last week we received a flyer from a realtor whom I'll call "John Doe". John obviously is pretty good at selling homes- after all, he's been in business for over 10 years with the same firm in the same office (we'll say he works for "Acme"). So, with a proven track record of solid business acumen and a professionally written and laid out color flyer, why would I do business with him?

I wouldn't. Why? He's got a winning smile, a good track record, and a really nice flyer.

Why? Because of his amateur, second-string personal email address- johndoe477@hotmail.com.

R U pulling my leg?

I'm sorry, but this is one of those things that just get me going. This IS the 21st Century, after all. And unless you're the only game in town, I'm not going to contact someone who has a throwaway email address- that's no way to run a business. So, my advice to the SOHO and SMB readers who may see this is this: get a domain that has your name in it. Give everyone in the firm a standardized email address, like john.doe@acme.com or jane.doe@acme.com. If your web presence is lame or nonexistent, hire a designer. Trust me on the website- believe it or not, I have "webmaster" in my list of credentials.

(Pause for effect)

Most readers will not know me personally, but on my best of days, I have trouble even spelling HTML. I'm kidding, of course. Just a bit. For example, if I see the tag <!DOCTYPE html> at the top of a page, I know I'm looking at an HTML 5 page. I also know how to recognize and change colors, as well as edit text and import or delete images. I told you that so I could tell you this: don't hire me or anyone like me to manage your online presence. There are pros out there who will do it better and for a fair price. If your business is your livelihood, you owe it to yourself.

But please do it.

Speaking of addresses, data and names, I'd like to share one more pet peeve. I don't consider pride to be a major fault of mine, but when it comes to business, please take the time to spell my name correctly... and get the city I live in right. The city I live in is a suburb of Chicago, with a population of approximately 60,000. It has three ZIP codes, one of which is shared by a village that is 1/15 our size and is strongly mob connected. I cannot tell you how INSULTING it is to get a piece of mail addressed to me, with the proper ZIP code and +4, but with that other village's name on it. Mail like this goes DIRECTLY to the garbage.

I work in an industry very closely associated with direct mail; most of our business is from folks doing direct mail. Every day, I look at mail data, and am constantly reminded of the old programming axiom GIGO (garbage in, garbage out). However, from a personal perspective, I tend to scrutinize the mail we receive at home... and Jennifer has also taken an interest in it. The USPS (United States Postal Service) depends on direct mail, but does not often see fit to enforce its own regulations. I see so much direct mail that is so out of spec that it really IS junk mail, and yet the USPS just accepts the money for the postage rather than enforcing its regulations.

That's not my beef, though.

Even though we live in a 24*7 world, take the time to get stuff right. Get the easy (data) stuff right- for example, it's spelled Coeur D'Alene. not Coeur Dalene. It's not <frigging> rocket science, folks. Then, when you get the easy stuff right, spell my name correctly- at least get my gender right! My name is a little long, and in an address block is often truncated. How about just truncating it to a standard, shortened (nickname?) name rather than just arbitrarily chopping letters off?

Okay, I'm putting the soapbox away now.

Data: 42 (my personal database project) is inching along. I got some data entry done over the last week, and did a few table upgrades. I think I'm currently up to seven complete CD entries, and quite honestly the last one was a bear!

I'm not a huge Elvis Presley fan, but I picked up a SEALED copy of his 30 #1 Hits CD from our local Goodwill store a few months ago for a buck (1USD). When I first built the table for CDs, I had included seventeen track titles. I knew I'd need more, as classical music tends to have LOTS of tracks. The Elvis CD, though, has thirty tracks, so I added thirteen tracks to the table, only to realize that there is one unnumbered track at the end! In unrelated database news, I added 885 pieces to my Lego collection, courtesy of the Amazon Prime sale on the 12th. As I was thinking about Lego and the database, I decided to only include Lego part numbers for parts that I actually possess.

That's it for now, fellow datacampers. As always, I am hochspeyer, blogging data analysis and management so you don't have to.

Thursday, May 26, 2016

Dude, you're getting a Dell, or, Resistance is futile.

We don't have many Dell computers in our home. Longtime readers may nod in agreement at this statement, but for the benefit of the uninitiated, we have many computers in our home. And while not all are in service, all ARE serviceable; that is, with just a bit of TLC, I believe we could have no fewer than thirteen PCs online simultaneously.

Granted, I would NOT do this for a number of reasons.

Firstly, most of these boxes were "designed for Windows XP". Now, I have nothing against XP- I have one machine that is an XP machine- and it's physically OFFLINE.

Next, most do not have green power supplies. Although all of the monitors in the SUL are now energy-efficient flat panels, the power supplies aren't really worth upgrading, as most of the PCs that are not in service are pretty much waiting to be cannibalized.

Finally, it's just not practical. Four individuals running 13 PCs in a SOHO environment?

The reason Dell came up is that a certain Dell PC crashed one of my external HDDs recently. I know for certain that it's a particular Dell because every time I plug the external drive into this Dell, it asks if I want to repair the drive. Mind you, this drive normally is hooked up to another Dell or a Lenovo or an HP or Compaq machine with no problems. This particular Dell has issues, though, and the last time I use the drive, I ignored the repair message, and in turn it made the drive unusable. I don't know why or how, but this machine is now persona non grata.

End of story.

Lego data news-

I'd been using Peeron's database for some time, but recently it came to my attention that BrickLink has a more complete database. Not only does it seem more complete, but the part numbers are properly formatted, they are already in a separate column and the descriptions do not succinct and do not seem to require editing. The data will still require a bit of formatting to fit my particular usage. That's the good news.

The bad news is that unlike Peeron's single page text file format, BrickLink displays its data on webpages. Not bad, as its still copy-and-paste. What's really from my perspective is that each page displays fifty part numbers.

There are 837 pages... a relatively small price to pay for superior data.

As always, I am hochspeyer, blogging data analysis and management so you don't have to.

Monday, May 23, 2016

Men Without Hats, and Data.

"In other data news, I've got what appears to be a workable solution for my internal Lego part number. It's fairly lengthy at seventeen characters, and from all appearances, this should be sufficient. I've begun the data entry on this, and tried a few trial sorts. So far, everything looks good, and this is officially stage 2 of the Peeron normalization."

As noted over a year ago in this space, No plan of battle survives first contact with the enemy, the best plans often don't get past the first round or two of testing. It's unfortunately true, and I'm referring here to the Lego portion of Forty-Two. I'm sitting on a challenge right now that's more conundrum than impasse. 

With only sixty-five entries, I ran into one of the proverbial "straw that broke the camel's back" records. My format had been "ANNNNNNNN.XXXXXXX", where A=alpha character, N=numeric character, the "."is a placeholder, and X=alpha or numeric character. It's large and unwieldy, but that doesn't matter, as I only need it to force order onto the Peeron data. So, I need to add a few more characters to the right of the "." placeholder. Still, I hate to redo stuff I've already done. So now, it appears that I'm up to nineteen bytes for the primary key. Hopefully, it doesn't grow more than that.

Before I forget, the Men Without Hats reference is, of course, to The Safety Dance, which I was humming when it looked like nineteen would, indeed, be a safe number. I just skimmed the entire list, and I'm fairly certain I can stick with nineteen.

As far as the rest of the database is concerned, I've been pounding away at my least favorite activity- data entry. It's still a really small database, and not even relational as yet. Three tables, containing in total 753 records.

I think that's a wrap. The internal issue I always have when doing these types of updates is keeping this more blog than change log. Oh well, speeds and feeds are gonna be speeds and feeds....

Lastly, a big shout out to my friends in Russia and Portugal- I hope you enjoy reading this as much as I enjoy writing it. 

As always, I am hochspeyer, blogging data analysis and management so you don't have to.




Friday, May 20, 2016

Richard Strauss, and Data.

This past week, I've managed to get in to work fairly early almost every day so far (Monday- Thursday); I think Wednesday was the only day I came in later, and that was intentional. I've also been getting up earlier, and the time before work has been used to work on the database. I've made some good progress on the few boxed sets of audio CDs that we own, which is where Richard Strauss comes into the picture.

Before I go any further, I should give you a bit of my musical backstory, kind reader. You may have surmised from previous posts (especially the most recent one, Axl Rose, and Data,) that I'm more Rocker than Opera-Goer. This is 100% correct. However, my musical tastes are fairly diverse... Yo-Yo Ma, The Beatles, Chevelle, C.W. McCall, Air Supply, Deep Purple, Newsboys, Sibelius, 80's hair bands, choral, etc. I firmly believe that any music that is good should be played loud when possible.  I like some blues, classic Motown, and a smattering of jazz. I'm okay with the various forms of trance, and dance music- if it's something that catches my ear. I can even deal with disco these days. Things I pretty much have no interest in are rap, hip hop, opera and death metal.

So, what's with Strauss?

Well, the album pictured above is a boxed set I picked up at a library sale for a solitary U.S. dollar. Cheap-value-SCORE! It's a three disc set, with a booklet nearly as thick as the CD case. The 330 page "booklet" has all sorts of details, not merely about this opera, but about this particular recording, its cast and conductor. It also has the complete text and lyrics of the opera in French, English and German. When I purchased it, I had decided that the weight of the boxed set alone made it a good buy... little did I know that this was a "reference" recording, one by which all others are to be judged. Its also conducted by Herbert von Karajan, a legendary conductor.

Still, what's with Strauss?

Data entry, pure and simple. As this past week has seen a renaissance of Forty-Two, I decided to tackled boxed CD sets. The problem is that this particular set has sixty-two tracks, all of which have German titles. I'm slow enough at data entry without having to import special characters, so I did what any reasonable human being would do: I looked up the recording on Amazon, copied the track list and pasted it into Excel. From there, I copied and pasted each track into Access. That's where I stopped with music- I still have three boxed sets to go, and then it's on to albums.

In other data news, I've got what appears to be a workable solution for my internal Lego part number. It's fairly lengthy at seventeen characters, and from all appearances, this should be sufficient. I've begun the data entry on this, and tried a few trial sorts. So far, everything looks good, and this is officially stage 2 of the Peeron normalization. I still need to add dimensions and clean up the text descriptions before importing it into a table.

As always, I am hochspeyer, blogging data analysis and management so you don't have to.

Sunday, May 15, 2016

Axl Rose, and Data.

It's been busy at work lately. And the rain has been frequent. And... it's Springtime in my little corner of the world. Consequently, our lawn got a "little" out of control.

(Spoiler alert: this post contains references to Guns 'N' Roses and their songs... 
and a few other musical references!)

Jennifer was kind enough to help out and mow the front lawn, which is what everyone passing our home sees. The back... well, that's another story- and my bailiwick.

My time had arrived. It had not rained for over a day. I looked out the back window and saw the grass blowing like Dust in the Wind. As my gaze landed upon the compost bins, I had to rub my eyes because I thought I saw a black tophat sitting on top of one of the compost bins, and a bandana on top of the other. I shook my head, blinked and looked again. Without warning, the yard had gone from broad daylight to a starless night. The yard's verdancy had turned to a monochrome with rough-cut video quality- I kid you not!. Our neighbor's white fence had been replaced by a gaggle of Marshall stacks, spotlights were illuminating the compost, and Slash, with trademark shades, ciggie and Les Paul, and partially obscured by the output of several smoke machines was furiously spewing out riff after riff, while Axl crooned, "You're in the jungle, baby" as he seemed to float over the grass. I covered my eyes, and shook my head. When I looked again, the moment was over. My yard was green, and back.

Seven Marshall stacks- fourteen cabinets and seven heads 
 Sheesh! Less Monster, more cowbell... maybe.

So, I cut the grass. Yes, it was long- long overdue for cutting. I'm not certain how long it took to complete the task, but I had to make several trips to the compost to empty the mower's bag. When that was done, I decided to tackle another yard task: the ivy. And, to be fair to the myriad of horticulturalists, botanists and assorted green thumbs in my audience, I'm not really certain what this plant is. It has dark green leaves, a woody stalk, and will sprout roots along the stalk. It's also quite capable of climbing. So, we've decide that it must go.

I got most of it off of the red mulberry. Now, technically the mulberry should be a bush, as it has multiple trunks. This one, though, is ~30' (over 9 meters) tall, so I call it a tree. Many websites also call it a tree.

Our Mulberry last Winter

I put the mower back in the garage, and grabbed a pair of leather gloves, a few yard waste bags and some clippers, and proceeded to tear into the ivy. For reference, the bags are constructed of a double walled heavy brown paper, and are approximately 16" x 12" x 35" (40.6cm x 30.5cm x 88.9cm). I filled five of them, and I'm probably only about one-third of the way done. This was sometime last week, and since then, we've had rain at least four times, so the next time I tackle this project, the stubble from the first round should be visible, which should make the next phase a bit easier.

Finally- data! (Sorry, no data pictures!)

On Saturday morning (May 14th), I FINALLY completed the 1st phase of the Lego Peeron data normalization! Speeds and feeds are appropriate here, so here we go!

The original Peeron database I'm using is from March 19, 2012. It contains 18,510 parts (rows). After the first passes through the data- in which the data started out as a text file, and then was converted to an .xlxs file, and then the data had its initial cleansing where a "base" part number was created- the data ended up totaling 16,218 parts. It should be noted that this not a fixed number, as there are more parts that need to go to the "Stickers etc" worksheet. I also want to emphasize that I am cleaning data and not normalizing; after all, all of this data still resides in an Excel file, so it is still a flat file. My next task is to impose order onto this file, so that will require the creation of a unique part number. This part number will be the primary key once it is imported into Access.  

Before any of this happens, I need to come up with a standardized format for my unique part number- currently called the "DB_Tracking_Number". Ugly, but good enough for now. I can't really automate this, once again because of the lack of standardization in the Peeron data, and my attempt to impose my own personal spin on Lego organization.

So, for now, I have ~ 16,000+ part numbers to create.

As always, I am hochspeyer, blogging data analysis and management so you don't have to.



Monday, March 14, 2016

Meanwhile, back in the Secret Underground Lair....

Those kind folk who have been following this blog for a bit may be aware that our home office, the Secret Underground Lair (SUL), is undergoing a bit of a reconfiguration. That's actually a "bit" of an understatement.

Before going to work last Monday evening (3/7 or 7/3 depending on your calendar), Mr. T and I spent about a hour of quality time grunting, straining and sweating as we resituated eighteen cartons of books. While moving the books, I discovered a few things: not all of the boxes or plastic storage containers were properly or efficiently packed, and more importantly, there are MANY books in storage that I need access to, so we will be swapping some books out.of storage and replacing them with books which really should be in storage.  It was worth it, though: for the moment, we have TWO walkways in the office. Mr. T tried some scenarios in Microsoft Visio, though, and is currently of the opinion that we will not be able to maintain this corridor if we want the best use of our limited space. I hope to do a bit of work in the Dungeon to free up some additional space. As we are on the cusp of Daylight Saving Time (ugh!), it's nearly midnight, so any more SUL work will take place tomorrow (Sunday 3/13 or 13/3!). And if anyone is doing the math, it has taken me seven days to get this far with this latest blog entry... overtime is great for our bottom line, but it definitely puts a damper on other things- like blogging!

Data!

In my previous post Working with someone else's data, I mentioned that I'm using peeron.com's data as the basis for my Lego data. I also indicated that although it's a great list, I have a good amount of work to do before it is useful to me. The following two lines are a sampling of consecutive rows copied from the .txt file, and illustrate quite well my "issues" with their data:

3001px2 Brick 2 x 4 with Eyes and Wavey Mouth Pattern   nose, nostril, snout
3001px20        Brick 2 x 4 with Red Danger Stripes and Two Horizontal Stripes Pattern

For starters, peeron.com has done a nice job of making a list of Lego parts. Their website is excellent, and their search engine is very nice. However, I'm not interested in their frontend, or even their instruction scans.

I'm cataloging and categorizing MY data, and their part number list is what I need. And although I've also mentioned before that this list is at LEAST four years out of date, up to date is not what I initially need. I need the basics first. And although peeron's 18,511 rows of data may seem impressive to someone who isn't an Adult Fan of Lego (AFOL), for me the list is horribly incomplete as I have purchased sets that were manufactured after 2012. Also, it has a lot of data that is extraneous for me, primarily part numbers for sticker sheets and parts which have somehow been redesignated, renumbered or somehow made obsolete. So, I need to do a bit of rework on their data.

The red bolded text above is two consecutive rows of data copied from their .txt file. Based on these two lines, I hope to show how much work I have cut out for me.

For starters, I took their .txt file and copied it into an Excel spreadsheet. My first task is to make two columns out of this data- one for the part number, and one for the description. I'm using 3001- this is the Lego designation for the ubiquitous 2x4 brick which nearly every human being has stepped on at least three times. The next few characters describe the variation/description of the brick. So far, so good. First problem: I don't know of an easy way to separate the part numbers from the descriptions- I'm certain there's some software that can parse the data, find the first empty space, and move the next set of characters into a new column. So, until I find a software solution, I'm doing CRTL+X, CRTL+V to manipulate things. Second problem: Once they are separated, I can't sort these, because although px2 comes before px20, px20-23 come before px3 in Excel. And some of these numbers run into the hundreds. So, nearly every suffix will need to be reformatted to actually display in numeric order. To do this, I'm going to create a new part number that my database will use; if I ever create my own Lego website, only the original Lego part numbers will be used.

After the columns are created (and before I start tweaking the part numbers), I'm going to pull out the extraneous part numbers and put them in a separate worksheet. After that, the remaining data descriptions will get trimmed down and standardized with a few hits with Find and Replace.

I've currently got 2,022 rows separated- roughly 11% of the 1st stage is complete. That's it for now.

As always, I am hochspeyer, blogging data analysis and management so you don't have to.






Saturday, March 5, 2016

Working with someone else's data

At a certain point in one's data career (plain English: pretty near every day, and with nearly every piece of data one normally touches!), you will have the "opportunity" to work with data someone else has prepared, or at least accumulated or aggregated or in some way modified. As I create Forty-Two, "most" of the data I will be using is data which I have entered (created) into a table. There is one notable exception to this, though, and it is for me a moderately to severely painful one: Lego parts.

As mentioned in a previous post, my experience tells me that the best place for my Lego data is in a spreadsheet- at least initially. And although calculations can be done in Access, Excel is much easier for me. And, they can be linked at a later date.

However, there is a HUGE caveat to the Lego data. The title refers to someone else's data. In my case, I'm using peeron.com's parts list. It's the best one that I've found, and many other AFOL (Adult Fan of Lego) sites use it.as a parts reference. The problem with the list and site is that they're horribly out of date. The list has a very standard naming convention, and the last update was nearly four years ago at the time of this writing. However, as it is the best, easiest to use and most complete list currently available, I've decided to use it. I can update newer parts as I find them on other websites.

The other issue with the Peeron data is that it is a .txt file. Not bad when importing to Excel- just copy and paste. However, to make it usable, I have to manually edit the 18K+ rows of data. As I'm in no great hurry to finish this phase of the project, it's not too much of an issue. Still, .... I know a guy (as they say) who might be able to help. More on that later.

The final issue with Peeron is that its creator strove to make it a very complete parts list. As such, there's a great deal of data which I actually don't need- stickers and superseded part numbers are two types which come to mind immediately. So, if my guy can fix my primary issue, the other issues will be much easier to deal with. If not, then I still have a great deal of work ahead of me. UPDATE: I'm not really surprised- but also not horribly disappointed- that we were not able to fix the data.

Before I forget, I wanted to post a brief update on the blog itself. I'm not quite sure when I did this last, but I think it was around the time the blog hit 10K viewers. As of today, the blog has over 15K viewers in fifty-five countries on six continents (c'mon, Antarctica!). Africa is represented by three countries, Asia by seventeeen, Australia by one, Europe by twenty-seven (I'm counting Russia in the Europe column rather than Asia), North America by six and South America by one. What's most amazing to me is that of these fifty-five countries, I could only count seven where English was either the official language, a dominant language, or one of a group of commonly accepted languages.

To each and every reader- THANK YOU!

Last: a small compilation of my blogs dealing with data (for those who are interested in how a small-time operator handles data)

http://hochspeyer.blogspot.com/2016/02/you-said-this-was-about-data-analysis.html
http://hochspeyer.blogspot.com/2015/06/data-science-pt-1.html
http://hochspeyer.blogspot.com/2016/02/a-database-against-rules.html
http://hochspeyer.blogspot.com/2015/05/data-defined.html
http://hochspeyer.blogspot.com/2015/02/forty-two-v7-or-so.html

As always, I am hochspeyer, blogging data analysis and management so you don't have to.



Sunday, February 28, 2016

Managing data better

I just discovered a fairly large data loss. As mentioned in a previous post, I rebuilt my laptop due to a failed HDD. Unfortunately (no one is looking, you can raise your hand if this has ever happened to you), I had a fair amount of data on that drive which was not backed up.

To paraphrase Hall and Oates, its gone. So, to paraphrase the unofficial motto of Chicago ("Vote early, vote often"), Save early, save often.

As promised in my previous blog, I'm going to spend a bit of time on Lego, because as far as my data is concerned, it is handled a bit differently than the other items which I am cataloging.

Lego and I go way back, but I didn't attempt to catalog my Lego collection until fairly recently. I've given consideration to including Lego in the "big database" (Forty-Two), but have found that- at least for me in this particular application, Excel is the better tool for me to use. Let me try to unpack that a bit.

Some time ago, I had the opportunity to observe how several different businesses utilized software to work with data. Some preferred Microsoft Access, and most preferred Microsoft Excel, or, to be a bit more generic, spreadsheets were preferred over databases. In only one business were both used- but independently, rather than in a complimentary fashion. The preference of software had little to do with function, generally speaking. Rather, it was more about culture and familiarity. In every case I observed, the results could have been improved simply by not just "thinking outside the box," but merely by thinking.

In my case, a spreadsheet seems to be the best solution- based upon my experience. My primary reasoning behind this is because Lego is a single thing which does not need to be linked (or, related) to anything else. And, even though I may be interested in a bit of analysis of the Lego "population", in the larger scheme of things Lego exists in its own unique bubble: shapes, colors, themes, sizes. I could build tables based upon these and other categories, but once again, they would relate only to Lego.

Forty-Two, on the other hand, illustrates quite well the differences between a flat database (the Lego spreadsheet) and a relational database (Forty-Two). With the Lego spreadsheet, I can see a snapshot of all the facts of my Lego collection. But with Forty-Two I can write an ad hoc query to tell me where all Lord of the Rings media is located. It would show where each book, soundtrack, game and video is located. It would also show the number of copies in a given format. It would also tell me the last time the media was viewed.

There's also one additional aspect of Lego databasing which is unrelated to software, but makes cataloging it so much easier: storage.

I'm an AFOL (Adult Fan of Lego). AFOL is a title; more of a descriptor, actually, as it really doesn't carry any of the clout that, say, CCNA carries. Still, it differentiates me from most adults who play with their kids while playing with Lego. And, its somewhat hard to say that without sounding like some sort of pompous jerk, because it sounds like I'm slamming parents who "play" Lego with their kids. Quite the contrary! If you're a parent who engages with your kids over a pile of Lego, kudos to you!

An AFOL, though, uses Lego as their primary creative medium. There are professional Lego artists out there who make a good living by building amazing models for corporate clients. There are also educators who use Lego in either a standard classroom setting, or in a program for ASD (autism spectrum disorder) kids. Even architects use Lego for models. And although each of these examples is an example of adults working with Lego for a living, the typical AFOL is something else.

They're a bit of a subject expert on Lego. They probably have some advanced building skills. But mostly, I think, they are something of a type of a Leonardo da Vinci. They are probably the precursors of the maker movement, which is another subject entirely.

Back on track: AFOLs need organization, and the English word for this is storage. Over the course of many years, many bricks (elements) will be collected. I've found that the (current) best way to keep track of these is to put them in plastic bags (connected in groups of 10) with a card inside indicating the count and the date of the count. These, in turn, reside in drawer organizers.   

As always, I am hochspeyer, blogging data analysis ad management so you don't have to.



Saturday, February 27, 2016

Feeds, Needs and Speeds

Well, it's starting to get real. The database, that is. As I posted earlier on Twitter, 200 rows of data in a single column do not a RDBMS make, but it's a start (the count is currently 224). As Jackie Gleason quipped in one of his signature lines, "And away we go!"

The casual reader may be wondering at this point: there are actually folks out there still building relational databases? Isn't the non-relational NoSQL model more popular in terms of new deployments, versatility and just plain coolness?

I'd like to explain a bit of my personal journey that brought me to Forty-Two, the database.

Long ago, like teenagers everywhere, I faced highschool graduation without a plan. Not just a clear plan, mind you- NO PLAN. Computer Science was in its infancy at the time, and nearly nonexistent in most high schools. I was not one of the cool kids, nor was I a jock or a brain, but I was also not a nerd. However, I knew one or two nerds. And the nerds were highly focused in their scholarly discipline: they were not into math and science; they were into math. They were geometry slingers, wearing leather slide rule holsters on their hips. They had glasses, thick glasses. Below average complexion. Few social skills. Pocket protectors in their left shirt pockets. And behind the pocket protectors... punch cards for their next "program".They were the late 70's analogues of Drs. Sheldon Cooper and Leonard Hofstadler (The Big Bang Theory).

I was not them. Well, sort of not like them. I liked history. I was (probably) the worst kind of history nerd: I was a military history buff. I started out with Avalon Hill games like France 1940 and Panzer Blitz, and progressed to Tobruk and Squad Leader, eventually culminating in the non-Avalon Hill classic Fletcher Pratt's Naval Wargame. The point is this: as the complexity increased, the playability decreased, as did the number of folks willing to take on the rules. But, I digress.

I was accepted to Rosary College (later Dominican University). I declared history as my major, and spent two unremarkable years there. I eventually had five colleges or universities under my belt, with no undergraduate degree to show for all of the buckazoids invested.

Long before this became a mainstream theory in education, I discovered that we do not all learn in the same way, and that higher education was not really the best choice for me.

Fast forward a few decades. I've mentioned this fairly recently- I was working as a data analyst at an "action sports" company. The company designed paintball equipment, and had it manufactured overseas. My job as the data analyst was to take Wal Mart RetailLink data and dice and slice it into what my employer could use.

The problem was my employer was using either Office 2000 or 2003, which limited Excel to a maximum of ~64K  (I believe it was 63,536) rows. As time went on, my data often exceeded this artificial limitation, and I was forced to use Access just to grab the Monday morning data. Once again, skipping several steps, I became adept at moving data between Excel and Access.

Fast forward one more time to today. I use all sorts of tools to do my job; Access and Excel aren't really part of my professional portfolio of commonly used programs on the job, but I use them at home,.. pause for effect.

Yes, although I may have mentioned a bit about Forty-Two before, I don't think I've said too much beyond it was my own little database dev world. Here's where the title comes in: way back when I was an I.T. reseller, we used to often qualify sales by talking about speeds and feeds- equipment specifications. Needs are also important (besides completing the alliterative trilogy).

So, Forty-Two is an obvious reference to Douglas Adams works, and is so named because its' goal is to answer that elusive question: what is the meaning of Life, the Universe, and Everything. It is starting out life as an Access database, currently with only one table. Previous iterations have taught me to take it easy with adding tables, so my aim is to get this Titles table to be mostly complete before adding additional single- or very few-column tables and then finally starting to create the relationships.The first table is called Titles simply because it holds titles: books, videos, software... if its media, then its name goes here. Why? Forced normalization: why go through the normalization process when I can start out with a relatively clean dataset?

This is starting to turn into a wall of words (by my standards, anyway!), so stay tuned... Lego is next!

As always, I am hochspeyer, blogging data analysis and management so you don't have to.

Monday, October 12, 2015

Geomorphology, meterology and Big Data

Edit: I found a glaring quantitative error, and corrected it.


I'd like to open with an apology to my long-suffering and patient loyal readers.

To be polite, my recent writing has been sparse at best. The reason is that I've been working a lot- I haven't had a "real" day off in three weeks. I think- this is my third weekend without a whole day off. I took Saturday evening off and will be back in the office on Sunday. And what, one might ask, does the author do for a living that requires so much work?

I'm a programmer, working in direct mail (a.k.a. "junk mail"). And what causes overtime in this field?

To keep it simple, there are really only two factors: workload and workforce. Here in the United States of America, the month of October is the time for senior citizens (folks who are 65 years of age or older) to make some choices regarding their prescription drug benefit provider. I'm not an expert on this, but the bottom line for my company is that in September and October we experience a huge spike in this business from these clients. This year, however, workforce came into play. There is another manufacturing plant that our plant has a fairly close relationship with, and they are currently shorthanded (as is our plant) and they also have a client that is doing a similar ad campaign. The deadlines have been tight, and everyone's resources- human and physical- have been severely stressed. My role in this has been pretty much support- but it's been for both plants. Or, wL > wF.

Having said that, I wanted to take a bit of a light-hearted look at Big Data.

There's been a few topics that I've been wanting to write about, but some recent twitter activity led me here. Just for the record, I'm currently listening to Steely Dan's "Midnight Cruiser", and thinking about data.

Data. Steely Dan. Yeah, not much of a connection there.

I'm not sure if the average reader realizes that data geeks even enjoy music.  To be blunt, we do.

But... back to data. Geomorphology is a real word. I was introduced to this term by my wife, who has a geology degree. Geololgy one- liner: she has rocks in her head. She said so. Anyway,  I find the big data landscape falling somewhat messily onto this collision of mismatched terminologies.

As I am not a true "data" person, I often laugh at data terminology and enjoy extending it to its ridiculous, but plausible limits limits.

Point:"data lakes".

Everyone pretty much understands (more or less) what "big data" is. Pretty much like everyone understands what "crime" is. Or "pornography". Alles klar?

In other words, aside from I.T. insiders and those who follow big data, no one really knows what big data is- or how pivotal it can be.

So, I suppose this is a call to action: how do you define your data?

I do not have a lot of data, relatively speaking. "Relatively speaking", of course, is a HUGE qualifier.

When I think about my personal data, I think in terms of things that matter to me- in the "real world",  these things have little value. In the real world, I tend to generate lots of data which has no value to me personally. For example, I've been on twitter for around three and a half years, and in that time have posted nearly 2900 tweets.That sounds like a lot of tweeting, but in reality it's far less than three tweets per day. What would be interesting to me would be a breakdown of my top hashtags.

But, as usual, I have digressed.

The personal data that I track is only in a few categories. I use data to catalog stuff, for the most part: books, videos, music and Legos. I also keep a pedometer log.

Most- if not all- of this data is useless to pretty much anyone except me. But, here we get a peek into the actual application of data science IRL. All data is data, but of all that data, which is most relevant to you? Does Lego care how many 3001 blue elements I own? I think not. They probably do care, however, about my age, where I purchase Lego products, and how much I spend on Lego in a month or year.

This is truly the science and ART of data science. Much of what I tweet on the subject of data science and data analysis is somewhat technical, focused on languages, algorithms and "sciencey" stuff... but business and ethics are also huge, and seem to be marginalized.

"What is the greatest Rock 'N' Roll song of all time"? A valid question. Of course, it is a question that cannot be answered- at least, not with data. Usually, there seem to be three contenders: "Hey Jude" (The Beatles), "Stairway To Heaven" (Led Zepplin) and "Freebird" (Lynyrd Skynyrd).

Likewise, a data scientist must be in tune with business: what is your best product/service? Data science should not only answer that question, but give stakeholders the answers to the five great press questions: Who? What? When? Why? and How? When a data scientist returns valid, data-based answers that are clearly communicated to these questions, the stakeholder has a valid representation of their business based on science and art.

Sorry- I never got around to the humor of Big Data... maybe another time.

As always, I am hochspeyer, blogging data analysis and management so you don't have to.

Monday, September 21, 2015

Down and out on a Sunday.

My Arduino Uno is sitting on my desktop in the Secret Underground Lair. It mocks me, sitting there connected to the PC with a 1M USB cable, it's onboard LCD flashing amber every second per its programming. And I am responsible for the programming. Tomorrow, though, all of that changes. Tomorrow is the scheduled arrival of "official" Arduino bases, and another Arduino. Jennifer will be getting my current Arduino, and I will be getting the new Arduino. Why, one might wonder... why indeed?

Because, Mr. T's Arduino has an Atmel 328 chip in a DIP configuration, whilst my current one is in an SMT package. I'm hoping to be 100% compatible with his board as we journey through the programming adventure together.

So, Sunday....

The weekend officially began for me around 0600 on Saturday morning, as it usually does. I got a decent amount of sleep before Jennifer, Mr. T and I headed off to The Bridge Community Church later that afternoon.We came home and Jennifer made tamales. They were incredible- and they are gone.

Sunday rolled around. When I woke up, my head felt like a balloon. Sinuses were off the scale in mucus output. I felt (*bleh*). I felt so (*bleh*) that I drank copious amounts of water for around eight hours. Kenji, our black and white tuxedo snowshoe cat, hung out with me for the better part of the day (he likes TV and he likes human presence while watching or listening to TV). I watched three complete football games. Not only is this unheard of- it is unprecedented. I was a couch potato- this is also unheard of.

That's how crummy and run down I felt.

The good news is that most of that seems to have passed. The bad news is that the weekend turned out to pretty much be a complete wash. That is, all of my plans were for naught. I had planned on a good walk, cutting the grass, fixing Jennifer's computer, and some quality "alone" time with my spouse.

None of these things happened.

It's now almost 0200, and I'm feeling much better, but my weekend is essentially over. Monday promises to be an adventure- Kenji and Kaley have their first visit with the veterinarian. Joy. 



The Monday recap- almost time to get ready for work. The cats survived their visit to the vet. It turns out that Dr. Chris is also a huge Lego fan, so much Lego discussion occurred while the cats were being poked and prodded. The new Arduino arrived along with the bases. I mounted the base to my old board first, and gave it to Jennifer. Then, Mr. T and I mounted the other two. I tested my new board out, and the amber LED started blinking immediately... I guess Arduino tests boards out with the blink program. So, for me to test it, I would need another program. The serial monitor program is the perfect sketch (program) for testing the Arduino for a couple of reasons. First, it's short- only two lines. Secondly, it tests two way communication with the board: write the program, compile, and upload. Then, open the serial monitor, and you should have a short message... this is the Arduino equivalent of "Hello, world."

As always, I am hochspeyer, blogging data analysis and management so you don't have to.

Tuesday, May 26, 2015

Sometimes even Nightstalkers drink decaf

I don't make a habit of drinking anything decaf... I'm a Nightstalker, for Pete's sake! My normal routine is a cuppa tea before going to work, and then a second once I arrive at the office so that I can be calm whilst reading my email. Tonight, thought, is the tail end of a three day weekend, so I'm not a Nightstalker. I'm more closely related to Mr. Mom than anything,

Before I get to the Mr. Mom thing, though, I'd like to talk a bit about tea. Quite literally, before Jennifer and I got married, I was the quintessential comic book/cartoon knuckle-dragging Neanderthal tea brewer (those readers who are from areas where there is a strong tea heritage might want to skip this part, or have one of those inflight distress bags handy. In fact, you might need a trash can). Back in my bachelor days, my morning tea ritual went like this: get a shallow sauce pan and fill it with cold water. Bring the water to a rolling boil. Without reducing the heat, carefully drop in one Lipton tea bag. Continue boiling until you can see a brown ring which marks the original "full" level of the pan. Turn off, discard teabag, and pour into a cup. Add a teaspoon of sugar. What were the attributes of this tea? Well, most lava flows in Hawai'i were less viscous that this stuff, and NASA has black holes on record that emitted more light than this tea reflected.The "flavour", if one could describe this brew, was somewhere between "turmoil" and "despair".

I'm not certain of what exactly it was that caused my tea preparation habits to change, but when I got married I went from Neanderthal to tea snob. My tea tastes have broadened quite a bit, although I generally still drink mostly black teas from the likes of Lyod, Tetley and Thompson's- as well as one or two Indian brands. For the most part, though, all of them are prepared in a similar fashion. Prepare cup by adding a teaspoon of sugar then the tea. In the case of bagged tea, it goes in the cup before the water. For loose tea, I have a red tea filter that is very close to the red that the Swiss company Bodum uses in their tea and coffee products. A rounded teaspoon of loose tea is placed on the filter, and when the water reaches a boil, it is poured into the cup- on top of the bag or through the filter. A timer, which has been set to one minute, is then turned on. When time has expired, the teabag is retrieved from the cup and given a gently squueze to coax the last of the amber liquor from the teabag; for the filtered tea, the filter handle is given a slight tap, and then it is allowed to drip a bit into the cup before being cleaned out. I have a slight variation to these procedures at the office. Although I have a filter and loose tea, I generally drink bagged tea- primarily because the coffee machine (which has the hot water dispenser) is on the opposite end of the office.So, in lieu of a timer, after the cup has been filled with water, I put a lid on it and walk back to my desk. With the pouring of the water, affixing the silicone lid, walking back to my desk and seating myself, approximately a minute has passed, so I remove the lid, give the bag a loving squeeze, and I have my beverage of choice.

So, decaf? Yessir, yessir, two cups full. Jennifer is still out of town, and while Mr. T has been quite good at emptying the dirty clothes hamper, he hasn't really bothered notifying me that the laundry is full. So, tonight I had three loads to wash and dry. In my defense, we've had a fair amount of rain the past few days, so I've had to wait for that to abate.

But I digress.

Some of you may be familiar with the saying, often attributed to Mark Twain or Benjamin Disraeli, but actually coming from an article by Leonard H. Courtney: "There are three kinds of lies: lies, damned lies and statistics." Truer words have probably been spoken, but as frequent readers may be aware, the unifying thread of this blog is data. And with that, ...

I received a letter from Commonwealth Edison, our electrical utility. Now, ComEd, like most if not all utilities, is trying to be "green". It's a sensible position to take, quite practical, and makes them look like good corporate citizens. As someone who is keenly interested in data and applications thereof, I read the letter with great interest. Now, I've received letters from them like this before, and they're quite interesting. A couple of graphs, some numeric comparisons, and some helpful suggestions on doing your part to conserve energy.

Judging by the graphs and numbers, we're energy pigs. According to their statistics, we used 84% more energy than our neighbors last summer.

We're just plain bad, right?       

Well, maybe. Maybe not. You see, our family is different from the other families in our neighborhood in a couple of significant ways which the lies- er, stats, cannot reflect. With few exceptions, there are two types of families in our neighborhood (and in our neighborhood, all of the homes are free standing, single family houses). One type is a young family with school-aged children (grades K-12), and the other type is retirees, including singles, couples, widows and widowers. Our family has four adults; the two adults who work outside the home have very nonstandard hours. Our older son works in retail, and his schedule can have him working any day of the week, sometimes getting up as early as 0530. I work nights, and usually get home around 0500, but often later. Because of the strange hours, my wife usually does not get to bed until 0100 at the earliest. So, during the week, our house might see four to five hours of "normal" nighttime. During the day, two or three individuals will be awake and active.

How about the other families? Mom/Dad get up at 0500-0600 to get the kids off to school and get themselves off to work. Dinner ~1800, bed for kids 2000-2300, bed for parents 2200-0000. These homes will have a "normal" nighttime of closer to seven hours.

Singles and retirees? Similar to the family hours, with retirees probably closer to eight or more hours of normal nighttime. Also, much less cooking, laundry and climate control.

All things considered, I don't think we're doing badly at all. In fact, given the additional information I've considered, there aren't really that many "efficient" neighbors that are actually efficient. Just one example for your consideration: in the past week, I think I've done five large loads of laundry; I'd bet that the widow down the street may have done one small load in the same period. Who's more efficient?

For truth in data, I've got a few speeds and feeds from my growing Lego database to share. From a development standpoint, it currently consists of four datasheets- I've done nothing so far with the fourth, as it is going to be the summary page. I'm fairly certain that I'll be adding a few more worksheets- what I currently have are basic bricks, plates and Technic. The current grand total of all elements (parts) is 6,475. In the For What It's Worth Department, I think the highest count I've ever gotten is 24,000.

Monday, May 18, 2015

Data, defined (part 2)

Right after I hit the "PUBLISH" button on my last blog, I realized that I wasn't done. I know I had the option at that point to pull the piece back and add the other thoughts, but I don't like to throw out a wall of words just because I'm not done... I'd much rather give the reader a break and come back another day, and so here we are today with a continuation of sorts, taking a closer look at microdata.

But first, an update from the home front.

Sunday the 17th was the third Sunday that Jennifer had spent in the Dallas area. Our older son was off at a convention, leaving Mr. T and I a very quiet weekend. That's a good thing, too, as I still managed to rack up a sleep deficit. I've mentioned a few times that I'm a programmer that works nontraditional hours. I refer to my band of coworkers and myself as Nightstalkers. The big plus and big drawback of being a Nightstalker is that one often gets to stay at work until the job is done, which can sometimes mean a fairly long day, but the plus is that we are compensated for that time. Saturday ended up being a late day for me- nearly eleven hours, and then a technician was coming over to the house for the Spring air conditioning checkup at noon. At some point before noon I decided that I could not stay awake, so I asked Mr. T to wake me up when the tech arrived. The tech arrived and did his thing. I wrote a check for his service, and then went back to bed, getting up some time around 2030. Looking back, I really don't remember too much of what I did except for a bit of work on the Lego database. I was back in bed ~0430, and up Sunday a little after 1230.

Sunday was warm and the humidity was palpable. I opted for some breathable training attire to cut the grass. I have to say that I am perfectly capable of wearing some pretty nice-looking clothing combos, but fashion has little place in my workout or working outdoors clothing choices. As it was both sunny and windy, I had an Aussie-inspired wide-brimmed hat with a chinstrap. The short sleeved shirt and shorts were both black sweat wicking workout attire, and the footware: orange sneakers. Blood orange red, actually. New Balance all terrain running shoes. Peer reviewed, double blind studies utilizing FLOOS and LRBL have verified that these shoes allow me to cut the grass 19.3% faster than the average suburbanite. You read it on the Internet- it's got to be true!

After cutting the grass, I figured I'd take a walk. One would think I'd have learned my lesson from the last time I did this (two weeks ago, actually). No. No I didn't. I grabbed a fanny pack (these workout shorts don't have pockets) and headed out. Approximately an hour later I walked back into the house, drenched in sweat carrying an empty half liter water bottle.

All of that is a great segue to microdata. Why? Well, for starters, I have an Omron pedometer. I have the option of publishing my workout data to their website- in which case, my data would be a part on Omron's small data, and quite possibly, fitness big data. My choice, though, is to upload the data to the Omron tracking program on my computer, making it MY microdata. In the FWIW category, I logged 6.2 miles (13.64km) today- my best day in nearly two months of tracking.

The Lego database is growing slowly. I'm using Excel 2007, and having to relearn some things. I'm sometimes asked what should someone learn in Excel to be useful on the job. Well, it depends on the job. Every place where I've used Excel I've needed at least a few things that no one else asked for- and none of these were financial or statistical environments (which tend to be a lot more predictable in terms of desired skills). The Lego counts stand as follows:  Basic bricks- 1 part number, 12 colors, 1666 elements. Plates- no counts as yet. Technic- 3 part numbers, 3 colors, 1049 elements. Total elements (pieces)- 2715.

As always, I am hochspeyer, blogging data analysis and management so you don't have to.

Wednesday, May 13, 2015

Data, defined

Although it is not my intent, I am certain that this post has the potential to step on a few toes, possibly bruise an ego or two, or ruffle some feathers. I may even get someone mad. Really e-mad.

For starters, I do not have any letters, diplomas, certifications and am not currently professionally employed in whatever one might consider the "data community". Whatever that might be. I do not claim to be an expert or have any special expertise or training in the areas of Big Data, the Internet of Things/Everything, Statistics, Analytics or The Cloud. I was once employed as a data analyst working with Small Data for a short time.

Whew!

So, who and what exactly am I?

I'm a guy who tweets (and retweets) primarily on the subjects of Big Data, IoT, programming and related topics. As far back as high school- maybe even earlier- I've been interested in data. It was either my music collection or Fletcher Pratt's Naval Wargame that gave me my start in classifying and quantifying. I remember even attempting to do a few music surveys way back when, and some of the respondents were unhappy because the polls were not simple popularity contests, but the answers were weighted based upon their position on the poll. Fast forward to today. I'm currently building a flat database of my Lego collection in Excel 2007 (why 2007? Because that's what I have on the computer nearest to the Legos!). This, in turn, will be added to my master database Forty-Two- so named because it answers the question of Life, the Universe and Everything.

Having said ALL of that, I'd like to start off by saying that the term "data" may not be as concrete as we are lead to believe. In my world, data comes in the following flavors: Big Data, Not-So-Big Data, Small Data, Micro Data, and Statistics. Depending upon the size of the dataset(s) and one's perspective, most- if not all- data can fit into more than one classification. Really? Sure. Case: say there's a hypothetical high school senior who is one of the stars of his basketball team. He's a good defender, doesn't get a great deal of fouls (below the league average), and is about average in scoring- except he leads the league in free throw percentage. Several colleges and universities are interested in him- they've got data on this fellow going back to 6th grade. That's data- to them. To me, a person who could care less about basketball- it's nothing more than a bunch of irrelevant stats. On the other hand, these same scouts would not be impressed by the number of PhD's that follow me on Twitter.

So, how big is a Big Data dataset? I asked a coworker. He wasn't sure, but thought a mail list might qualify. Don't laugh too soon- some of the mail lists I've seen have more than 10 million names. To me, though, I'd put that in the Not-So-Big Data or Small Data categories. The IoT,  Amazon, Google, Youtube and Wikipedia definitely fit into the Big Data category, but to the average person, these can be tough to visualize. So, for what I think might be a decent, understandable Big Data dataset, I propose the 2010 U.S. Census. It was a 10 item questionnaire (with a few extra answers possible) that mailed to 135,000,000 addresses representing approximately 309,000,000 persons.

Small Data could be a database, a website or the phone directory of a small to medium sized city- the lines are pretty fuzzy here.

Lastly, there's microdata. I'm not sure if this term is used anywhere else, but I find it to be a convenient term for personal data- data generated and maintained by one person or one family for their own use and not often formally shared. A cataloged collection of coins, stamps, recipes, exercise/workout logs or Legos- all of these are Microdata in my worldview.

Thanks for your patience- I hope you enjoyed this. I generally write a lot less... I'm not a fan of writing or reading walls of words!

As always, I am hochspeyer, blogging data analysis and management so you don't have to.

Saturday, May 9, 2015

Life in the Twittersphere

I'm not going to do the parody lyrics here, but bring up the Eagles' "Life In The Fast Lane"  from their "Hotel California" album and you can sort of hum along- it pretty much works. I used to do a lot of parody lyrics, as a matter of fact, but as far as this blog goes, I really didn't make the connection with the Eagles until I actually saw the title of the blog.

For what its worth, this was going to be another "Hoodie Migration" blog, but that was before I read a few posts from the Outmannedmommy blog. To be fair, the Outmannedmommy blog is quite funny, but it may be NSFW (not safe for work) as F-bombs occasionally detonate- especially from guest writers! Still, I enjoy Mary Widdicks' writing- because her style is great, and I can relate to her writing.

Also, strangely (or not?), I don't read a lot of other blogs, Mary's is the only funny one I normally read; most of the rest are retweets on Twitter (nothing wrong with that) about the IoT and bigdata.

And here's the crux, the nexus of this post: Me and Twitter (if you ant to try your hand at parodying this, try Tom T. Hall's most excellent "Me and Jesus" or "Faster Horses, Younger Women, Older Whiskey and More Money".

Twitter is by far the number one reason for my dearth of blogs as of late. I've been contributing a fair amount to the Twittersphere, focusing primarily on Big Data and the Internet of Things. One of my coworkers constantly tells me that I'm a fraud (he's kidding... mostly) because I'm not a PhD, a data scientist, a statistician, or any number of other things associated with Big Data and the IoT... and at that point I always remind him that I make no claims to be anything more than a catalyst or a conduit for ideas. And even though I'm no expert, I've become fairly adept at picking interesting and informative stories, factoids, infographs and blogs. This is evidenced by several folks that ARE PhD's, data scientists and statisticians who follow (and retweet) me.

That's the really short version of Life in the Twittersphere, but I need to get to bed, so I am going to close this and hopefully get another blog out very soon.

News from the Secret Underground Lair (SUL): I think I had mentioned that Mr. T and I were planning a bit of remodeling in the SUL. I'm happy to report that the majority of that work is now complete. Well, the furniture moving is done. Everything that got moved out now needs to be moved back in- I wish there was a defrag utility for small offices... press a button on the wall, and all of the books, boxes, doodads, computer components, computers, etc would find the optimal place to rest and automatically go there.

Finally, data news (this blog is about data, right?) I've got a spreadsheet roughed out, and have put a bit of test data in. So far, everything looks good: according to my flat database, I have 71 Lego elements. 

As always, I am hochspeyer, blogging data analysis and management so you don't have to.

Saturday, April 25, 2015

Data, evolving

Quick- fill in the blank: I do some of my best work in ______________________.

If you filled that blank with "the bathroom" congratulations! You, Freddie Mercury and I think alike. I believe Freddie Mercury claimed to have come up with the idea and melody for "Crazy Little Thing Called Love" in ten minutes while soaking in the bathtub in Műnchen (Munich), Germany, and the same thing happened to me Thursday evening!

Well, sort of... I didn't actually come up with a #1 U.S. Billboard chart single for a supergroup in a hotel in the Black Forest in Germany, but I think I've come up with the way I want to get my Lego data into Access.

A few years back, I had designed a flat database for my Lego collection in Microsoft Excel 2003. It was pretty nice and fit my needs at the time quite nicely. It was composed of several pages where the data was entered, and then a couple of pages which gave some bird's eye level analysis via links to the data pages. The project was scrapped mainly due to my impatience with data entry and my inability to actually get an accurate count of all of those elements- besides, they are much more fun to build with than count.

So, my plan is to go back to that format, but instead of merely having links within the workbook, have links from the Excel 2007 workbook to the Access 2007 database. And why 2007 rather than 2010 or 2013? Expediency: the Legos are close to a machine in the Secret Underground Lair (SUL) which runs 2007. The SUL redo is going very slowly, but if I am successful in liberating space for my laptop in my corner of the SUL, then it will be done on the 2013 versions of Excel and Access.

I'd like to get started on this first- creating the spreadsheet with the data, and then importing the data into an Access table and then (hopefully!) linking each data cell in Excel to its corresponding Access field so that the database can be updated via Excel. Its been some time since I did this, and it was on the "pre-ribbon" interface, so it looks like I'm going to be relearning some Excel tricks.

In other data news, Jennifer and I have been walking more as the weather has become a bit more pleasant. My replacement pedometer (identical to my expired Omron model) is chugging along, but its going to be at least another month before I have daily month-to-month data available.

As to my new glasses: the lady who fitted me for frames was slightly surprised as to how quickly I adjusted to my new spectacles, although I think I'm still breaking them in. She also said something that caught me off-guard: I wear my lenses lower than most folks.

 I've never worn gradient lenses before- I've worn reading glasses since the age of eighteen and I'm now 50+; my new specs have three prescriptions per lens: top for distance, middle for computer work and bottom for reading. Although the glasses really do a good job (especially while driving), they are singularly ineffective in the place where I need them the most: computer. Here's the deal: before I went in for the exam, I had made some careful observations and mental notes about my work environment and how I typically used glasses, as well as my concerns about night driving. The gradient lenses the optometrist came up with (pretty much six prescriptions- three per lens!) are fantastic for general purpose applications (life!). My hope in getting this type of lens would be avoiding a second pair of glasses. To quote "Mad Dog" Tannen, "You thought wrong, dude." (*BLAM*) I was very quickly able to pick up the reading area of the lenses, and the distance area (which I don't really need but seems to be beneficial in low light/nighttime conditions) also "snapped" into place pretty quickly.  I really need the computer part, though, and this is where I'm experiencing issues. Simply put: the lens cannot focus on the whole screen- even when I push the frame up tight to my face for the best focus. In the "good ole" 15" CRT days, these glasses would have been miraculous; with the 23" CRT's I have at work and at home, the lenses are just not up to it. And honestly, at home it really doesn't matter, because I'm generally not looking at the whole screen. At work, though- I do a lot of page layout and composition, and I generally NEED to take in the whole 23" diagonal screen at a glance- and more often than not, at a 90 degree offset. So, after only two days, I was back getting fitted for single vision computer glasses! (Just a hint here: when your livelihood depends upon a durable medical appliance- even if your employer or insurance doesn't pay for it: BUY IT!) The good news is that the second par of specs was massively discounted, and I should have them in time for my return from vacation.   

One final note before I call it a night: my twitter account @CjoelHarrison grew 10% in twenty-four hours!

As always, I an hochspeyer, blogging data analysis and management so you don't have to.