Showing posts with label Led Zepplin. Show all posts
Showing posts with label Led Zepplin. Show all posts

Thursday, December 24, 2015

More "Pause" than "End"

Some years back, on a particularly busy night, a new hire rubbed eyes and blearily asked the Nightstalker crew, "Just how long DO you work?"

I answered in the only possible way: "Until the job is done."

That, truly, is what makes the overnight hours in Data Services interesting. I say "overnight" because my hours can be a bit fuzzy- and this is true for everyone in our department. Some folks come in fairly early, some come in around the traditional office start time of 0800, and some come in at 2200 or even later. We have one person who works a traditional second shift, and the remainder of the crew work something of a hybrid 3rd shift, starting anywhere from 1800 to almost 2300. I don't believe there's anything like an official policy regarding this; it's more of a rule of thumb: as long as one works one's 40 hours, everyone is happy. But, back to "until the job is done."

I think I may have previously mentioned that our plant can be considered a "job shop" or "contract manufacturer." In other words, the product that we manufacture- direct mail- does not "belong" to us. You will typically not see our name on it anywhere. We may have some involvement in the design process, and we are often involved in the postal and logistics components of a job, but at the end of the day, we're simply the folks who manufacture something for someone else.

Sometimes...

Sometimes things just need to get done. One of our top clients has a job that runs every month, and it keeps the programmer pretty busy for around a week. However, once the client approves the job, yours truly ("me") typically swings into action. This particular job requires a lot of hard copy:  proofs for press and prepress, and tray tags for press. Printing, collating and marking up of this takes ~12 hours, and is best done when the office is empty. Why? Because one simply needs a lot of space to do the job correctly and efficiently, as well as tying up two printers.

So, the Friday before last, I came in with the expectation that I would be working on this job. I had put a dent in it when I left Saturday morning, and came back Saturday evening. I got home around midnight, and everything was done.

In a nutshell, my weekend was more of a pause than an actual end to the week. The obvious downside to weekends like this is that they are short. Jennifer and I got in a nice walk on the following Sunday- almost three miles. I did a bit of work in the Secret Underground Lair, as did Mr. T, and great progress has been made. My plan for the weekend, though, had been to get to the unused desk and hook up my Raspberry Pi. That hadn't happened yet, but we're on the cusp of it!

One other huge thing that happened the previous week was the Great Book Cull, during which time a few hundred books were donated to our local library. Although both Jennifer and I hate getting rid of actual paper books, there were a few compelling reasons for eliminating these particular volumes from our collection:

1) They were classics (Shakespeare, Tolstoi, etc)- these are readily available online and downloadable as plain text files

2) Cold War fiction- entertaining reading when it was written, but no longer plausible.

3) Out of date software texts- bye, bye, Windows 95 For Dummies, etc... I  kid you not.

So that was the two weekends ago. I visited the opthamologist for the last post surgery follow up for my right eye. (all is good- come back in a year). Don't get me wrong- if you're in the Des Plaines area, Dr. Winkler is the best. I'm just ready to get on with my life!

Pause, continued... this past weekend (Dec 19-20): I got out on time on Saturday morning for a change, but I knew I'd be coming back on Saturday evening. I figured a few hours would take care of a few quick production QC's- boy was I wrong.

The first job went smoothly, and it was done fairly quickly. The second, though, took much longer, as some paperwork was missing. It's not unusual for this to happen with this particular job, though. The third job was unexpected, but pretty easy- especially since the lead programmer was also in the office to explain a few things. The last job, though.

In a nutshell, it was programmed correctly but the output was set up incorrectly, and consequently compiled incorrectly, and the output was just plain wrong. As it was late and hadn't brought my meds to the office, I remade the files and recompiled the job, but did not perform the production QC.  

Finally, data.

I need to say something about my writing process, which will eventually relate to data. I promise.

When I write a blog, I generally have an idea of where I want it to go. Sometimes it goes there, sometimes it doesn't. This blog is a case of: IT DOESN'T CARE WHERE I WANT TO GO.

Point: my idea for this blog actually had to do with a concept for music opinion data analysis.

Way back when, in my college days, I started my music database. It was on looseleaf paper, and it was unwieldy and incredibly difficult to update or edit. At some point, I decided I'd create something of a personal "Top 40". Only looking at my stats was somewhat biased, I quickly surmised.  Undaunted, I expanded my pursuit to include what others were interested in. When I shared my results, however, some (many) of the participants were upset when I told them the results were weighted. That is, their top votes received 10 points, and their bottom votes received one point. Huh?

Participants in the survey were asked to vote for up to 10 of their favorite albums, and 10 of their favorite songs. I asked that they please put them in order of preference. When the results had been tabulated, I shared them. In my circle of friends and acquaintances, no one could believe that "Bread's Greatest Hits"  was more popular than "Barry Manilow Live" (joke).

I did it this way for a couple of (what I thought) were very good reasons. First: everyone probably has a favorite song. Hey, if its YOUR fave, why not weight it? Second, I wanted to be able to track which songs were actually more popular than which songs were most often voted for. Taking the two most "popular" songs of this era (from my polls), the hands-down winners were The Beatles' "Hey Jude" and Led Zepplin's "Stairway To Heaven". As my survey was quite small, let's say I had 10 voters: 4 went with the Beatles, and 4 with Led Zepplin. Let's say the other two went with Lynyrd Skynyrd's "Freebird" as their favorite. If Hey Jude == 1, and Stairway == 1 and Freebird == 1, then we have 4 to 4 to 2. What matters next is what is the next greatest song.  So, two of the "Beatles" guys vote for  "I Saw Her Standing There", and a couple of Zepplin fans vote for "Rock and Roll". After that, each of the other Beatles and Zepplin fans votes for Hey Jude and Stairway, furthering the tie. Then, we factor in the Skynyrd fans, who generally might like the Zep guitars, but can sing to Hey Jude. So, the final score ends up being Beatles: 5.8 (among Beatles fans), Zep: 5.8 (among Zepplin fans). The two Skynyrd fans, though, both vote Hey Jude #2, and Stairway #3, giving the Beatles a total of 6.6 and Led Zep 5.8. Therefore, in this tiny, theoretical sample, "Hey Jude" is the greatest rock and roll song of all time.

Sorry about the length of this... I told you all of that to tell you this: I have a new data project which is based on my previous data project.

I know what songs I like and those which I don't like. I plan to quantify this data and publish my findings in a new format.

As always, I am hochspeyer, blogging data analysis and management so you don't have to,

Monday, October 12, 2015

Geomorphology, meterology and Big Data

Edit: I found a glaring quantitative error, and corrected it.


I'd like to open with an apology to my long-suffering and patient loyal readers.

To be polite, my recent writing has been sparse at best. The reason is that I've been working a lot- I haven't had a "real" day off in three weeks. I think- this is my third weekend without a whole day off. I took Saturday evening off and will be back in the office on Sunday. And what, one might ask, does the author do for a living that requires so much work?

I'm a programmer, working in direct mail (a.k.a. "junk mail"). And what causes overtime in this field?

To keep it simple, there are really only two factors: workload and workforce. Here in the United States of America, the month of October is the time for senior citizens (folks who are 65 years of age or older) to make some choices regarding their prescription drug benefit provider. I'm not an expert on this, but the bottom line for my company is that in September and October we experience a huge spike in this business from these clients. This year, however, workforce came into play. There is another manufacturing plant that our plant has a fairly close relationship with, and they are currently shorthanded (as is our plant) and they also have a client that is doing a similar ad campaign. The deadlines have been tight, and everyone's resources- human and physical- have been severely stressed. My role in this has been pretty much support- but it's been for both plants. Or, wL > wF.

Having said that, I wanted to take a bit of a light-hearted look at Big Data.

There's been a few topics that I've been wanting to write about, but some recent twitter activity led me here. Just for the record, I'm currently listening to Steely Dan's "Midnight Cruiser", and thinking about data.

Data. Steely Dan. Yeah, not much of a connection there.

I'm not sure if the average reader realizes that data geeks even enjoy music.  To be blunt, we do.

But... back to data. Geomorphology is a real word. I was introduced to this term by my wife, who has a geology degree. Geololgy one- liner: she has rocks in her head. She said so. Anyway,  I find the big data landscape falling somewhat messily onto this collision of mismatched terminologies.

As I am not a true "data" person, I often laugh at data terminology and enjoy extending it to its ridiculous, but plausible limits limits.

Point:"data lakes".

Everyone pretty much understands (more or less) what "big data" is. Pretty much like everyone understands what "crime" is. Or "pornography". Alles klar?

In other words, aside from I.T. insiders and those who follow big data, no one really knows what big data is- or how pivotal it can be.

So, I suppose this is a call to action: how do you define your data?

I do not have a lot of data, relatively speaking. "Relatively speaking", of course, is a HUGE qualifier.

When I think about my personal data, I think in terms of things that matter to me- in the "real world",  these things have little value. In the real world, I tend to generate lots of data which has no value to me personally. For example, I've been on twitter for around three and a half years, and in that time have posted nearly 2900 tweets.That sounds like a lot of tweeting, but in reality it's far less than three tweets per day. What would be interesting to me would be a breakdown of my top hashtags.

But, as usual, I have digressed.

The personal data that I track is only in a few categories. I use data to catalog stuff, for the most part: books, videos, music and Legos. I also keep a pedometer log.

Most- if not all- of this data is useless to pretty much anyone except me. But, here we get a peek into the actual application of data science IRL. All data is data, but of all that data, which is most relevant to you? Does Lego care how many 3001 blue elements I own? I think not. They probably do care, however, about my age, where I purchase Lego products, and how much I spend on Lego in a month or year.

This is truly the science and ART of data science. Much of what I tweet on the subject of data science and data analysis is somewhat technical, focused on languages, algorithms and "sciencey" stuff... but business and ethics are also huge, and seem to be marginalized.

"What is the greatest Rock 'N' Roll song of all time"? A valid question. Of course, it is a question that cannot be answered- at least, not with data. Usually, there seem to be three contenders: "Hey Jude" (The Beatles), "Stairway To Heaven" (Led Zepplin) and "Freebird" (Lynyrd Skynyrd).

Likewise, a data scientist must be in tune with business: what is your best product/service? Data science should not only answer that question, but give stakeholders the answers to the five great press questions: Who? What? When? Why? and How? When a data scientist returns valid, data-based answers that are clearly communicated to these questions, the stakeholder has a valid representation of their business based on science and art.

Sorry- I never got around to the humor of Big Data... maybe another time.

As always, I am hochspeyer, blogging data analysis and management so you don't have to.