Thursday, October 3, 2013

The Gathering Storm in the PC Gaming Industry

As a result of recent events in the computing world, I think we are entering into a very special time in PC gaming. This industry has been close to my heart since I loaded up Police Quest 2 on my first turbo-button equipped and ultra-beige PC clone. Since then, there have been many evolutions of the gaming platform but no revolutions. That may finally change in the near future. Despite the decline in PC sales, the PC gaming industry continues to grow, and the PC as we know it is morphing into device types that don't fit the molds to which we're accustomed.



With this article I aim to illustrate why we're about to go through a particularly intriguing time in gaming by covering how we got here, where we are now, what has recently changed, and finally why the future is so exciting.

With that said, let's start with the foundation of the modern gaming platform; we'll cover a bit of history and the pace of innovation. 

Hardware

My new PC Gets 35 MPG. 

  • CPU Performance has increased steadily and closely followed Moore's law since the adoption of the Intel 8080 in the late 1970s. Intel has largely dominated this space from then until now save a few year interruption in the mid 2000s when AMD introduced the K8.
  • Gaming relevant display adapters started with "Windows Accelerator" boards (Late 1980s) and since the introduction of the first 3d accelerators (Mid 1990s), the progress has been steady with minor interruptions by things like programmable shaders.
  • Audio progression was initially rapid and has stagnated significantly in the last ten years. Things got kicked off in 1988 with the Roland LAPC-I, then the Adlib and then on to the Creative Labs cards which is where we remain today. We nearly had a major evolution in gaming sound in the late 90s with the introduction of A3d by Aureal, but unfortunately they were forced out of business by the legal maneuvering of Creative. 

Operating Systems

Everything Old is New Again.




For the popular gaming platform this one is easy. After the C64 and other early platforms faded away it's been Microsoft, Microsoft, Microsoft. DOS to Windows. DOS was new then it was old. Windows was new then it was old. NT was new now it's aging, and the commitment to the OS as a gaming platform has waned for years.

Input and Output Devices

Because the IBM Model M Still goes for Sixty Bucks on Ebay.

Little has changed with the mouse and keyboard in the last 30 years. As with everything else, performance and ability to customize has increased bit by bit. At least most gamepad supporting games these days have standardized on the Xbox360 controller layout rather than forcing you to map buttons to your gamepad du jour. 

The monitor hasn't progressed much since VGA was introduced in 1987. Slow and steady improvements, with a dip in performance when LCDs became the popular solution. As it pertains to pure gaming performance (input lag/pixel persistence/response time), LCDs are just now catching up to CRTs.

Evolution Summary

Putting it together.

As you can see, not much has really changed given we're talking about nearly 30 years. We're still playing similar games on similar platforms using similar interfaces. That's not a complaint mind you, but relative to the change we have seen in other sectors in recent years this space seems long in the tooth. Now let's shift to the present...




Setting the Stage for Change

It takes a nation of millions a few companies to hold us back.

First, a few facts: Depending on who/where/how you ask, the average gamer ranges from 30 to 37 years old. Nearly half are women, and over half of Americans play video games.



The gaming industry is an odd thing when you look at it as a whole; while we take the status quo for granted it really has not fully evolved. My generation (Gen X) is the first generation that has grown up with this entertainment medium, and while we've aged & settled down the market has not. There still exists a perception among the baby boomer generation that games are "kids stuff", and with boomers still in control of many large companies it's no surprise that these issues exist. A single viewing of the "Spike Video Game Awards" is all you need to illustrate the damage of this ill effect. This is improving with time, but like all new mediums before it, the gaming industry still needs time to realize its potential. We're still waiting for anything the likes of 2001: A Space Odyssey or The Godfather.

We're Getting Close Though

Couple this with extreme confusion on the part of the gaming system providers and the conflict of interest that it creates. With Microsoft being the predominant PC gaming platform one would think they would be eager to lead the space. I think their own conflict of interest and lack of a credible competitor has caused them to become complacent. Time and time again Microsoft has claimed to commit to the platform, but since the release of the original Xbox in 2001 they back down in favor of that lone consumer-image win. The most important and perfect example of this is DirectX, which started out strong in 1995 with ten major releases in the first ten years, but since 2006 has seen only one major release (DX11). To help put this in perspective, Internet Explorer 7 was released the same year as DirectX 10. Many developers in the industry have been complaining for years that DirectX is holding the industry back and it even appears at times that Microsoft uses relatively minor updates in an attempt to push gamers to the newest version of their operating system. (See Windows 8.1 and in a more defensible example, Windows Vista) Update 10/17/2013: Microsoft has published an article in defense of the Direct3d progress.

Additionally, Microsoft doesn't want you to hook up a PC to your TV. While home theater computing has never been a major market, it was inevitable that Microsoft would be a major player because of industry partnerships and the hardware support picture. Now they want to cut costs in that area and push the sale of the Xbox One instead, and that leaves space in this, albeit small, market segment.

And this is just the start. The desire to move away from the traditional PC platform has been an undercurrent for years, but without a viable competitor there is no reason for hardware manufacturers to spend money developing for an alternative. That's not to say others haven't tried though, and their frustration can be best summarized by this (NSFW, Language/Text) statement from Linus Torvalds.

Market Trends and Impact on Hardware Support

"Close the watertight doors." -Capt. EJ Smith 


Mobile Computing is Gaining in Popularity

With the monumental rise of portable computing (Smartphones, etc.) there is an effect here that is easily overlooked. Android (Linux) and iOS are absolutely dominating the mobile space, and with that comes the demand for and expertise in coding for other operating systems. This sort of experience doesn't directly undermine the current PC gaming picture, but it helps to build an environment more suitable for change.

Microsoft seems primed to slowly abandon the traditional PC space in favor of pursuing the built-for-purpose device model of Apple coupled with the software-as-a-service model of Google. It isn't that they want to abandon gaming, they just want you to use a device for that. Over the course of the last year and a half they have taken definitive steps to reorganize and re-purpose themselves as a devices and services company.

The problem with this strategy is they're trying to sell a platform that, when compared to the traditional PC or newer solutions (below), limits functionality. PC gaming thrives in part due to the community's ability to customize, add-on, and change the experience. As of now, nearly anything one may want to contribute, be it modding, culture, or even whole games is open more often than not. In a future Xbox (or PS4) gaming space like the one Microsoft has portrayed those things become more difficult or impossible to do for the average lone and unassociated contributor.

Enter the Contender

Is that the protagonist from Mike Tyson's Punchout?


As I'm sure you're aware, Valve recently announced the SteamMachines, Steam Controller, and most importantly (to my story at least),  SteamOS. This article isn't about an in-depth analysis of these announcements, but for the sake of context I'll introduce each:

SteamMachines/Controller

The SteamMachines aren't a platform but rather a series of similar platforms. It's more of a hardware spec for a living room focused PC. While Valve will be making (and offering in a limited 300 person beta) their own "SteamBox", the playing field is wide open for any manufacturer to create and ship a SteamMachine. Boy if Bissell ever wants to do a market sector transition this one would be easy. The controller is a stab at providing a traditional controller footprint that will more easily map to keyboard/mouse focused games. Initial impressions on the haptic feedback mechanism and other details have been relatively positive. That said, the controller is really an attempt to bridge the "couch gap" created by targeting this platform to the living-room/entertainment space. Should you choose you can still use an Xbox360 controller with the appropriate receiver or a traditional keyboard/mouse.

SteamOS

This piece is what I find the most interesting and has the largest potential for impact. The software that will be run on any SteamMachine, SteamOS, has the following characteristics:

  • An entire operating system distribution based on the open-source juggernaut Linux
  • Will run and be supported (to an extent) on nearly any standard PC platform
  • Capable of streaming games from your existing Windows gaming PC, which not only leverages any existing heavy duty hardware but also helps to bridge the Linux compatibility gap. 
Valve CEO Gabe Newell and His Heavy Friend
Since this platform is based on Linux it creates quite an opportunity to leverage community improvements to its core. Say you want an API to have a certain feature for an upcoming piece of hardware... the vendor, the community, and Valve could all work in tandem to provide the solution quickly and efficiently for all. If exercised correctly, the partnership between private sector and community development can be very powerful. All that's needed to facilitate that partnership is financial incentive mixed with a motivated, intelligent user base. Both those elements stand ready to contribute should the pieces fall into place.

Contender to Whom?

I want you to be nice until it's time to not be nice.

So this is where I'm supposed to make grand statements about head-to-head match-ups with consoles and/or PCs, but I won't. I don't think those match-ups will really happen. This platform isn't taking on any of the established players directly.... at least not yet.

The beauty of this idea is that it is potentially more sustainable than any other model. While the purely private entities need to fight over exclusives and drop lots of cash developing their platforms, the SteamOS initiative can leverage an existing and growing library of top-quality gaming experiences and could rely on the slow, steady pace of innovation of the underlying platform to gradually drain market share from its competitors.

Wildcards

Nobody gave the 2007 Giants a chance.

This is the particularly interesting part of this equation. Seeing as how the PC platform is arguably (in theory) the most adaptable gaming platform that is or will be available, it is worth tracking several new technologies that may contribute to the success or failure of the SteamOS or traditional Windows based gaming PC platform:

Oculus Rift and other Output Devices


Jimmy Fallon Models the Oculus Rift and Looks... Comfortable?

The Oculus Rift is the first VR headset that stands to be both cost effective and performant enough to catch on in the consumer space. This device is a potential slam dunk for the living room space (space to move around + reliance on controller). Whoever has the most effective support out of the gate could dominate a small market of enthusiasts and genre devotees (flight sims, racing sims) right off the bat.

New Input Controls



Not to be left behind, the input space also has some new VR initiatives. There are quite a few out there, but the most interesting I've seen are the relatively portable treadmill solution Omni, the Stem VR controller system, and the Tactical Haptics controller feedback system. Again, these are longshot devices that will only drive a small market initially, but the concept of a PC in the living room is a much better fit for any of these than a PC in the office or a console with limited support.

Rebirth of the Home Theater PC

By targeting the living room space the SteamMachine/SteamOS combo also taps into another potential market: the aforementioned HTPC crowd and additionally, if the price and functionality are right, the growing "cord cutters" crowd.

Should Valve's media playback efforts prove not only sufficient, but best in class, it would open up the market for low-spec/low-cost devices from hardware partners targeted explicitly at that space. The prospect of a roku-like device that can also play casual games and stream more power hungry ones could be a very interesting alternative to the bevy of devices found in many living rooms now, especially if a solution was provided to play games on low-end hardware. Speaking of transitions:

Network Streaming


"I'm not really here, I'm in a datacenter."

Valve has already announced that you'll be able to stream games from another PC on your local network, but what if they were to introduce a "cloud" gaming service down the line? The technology is feasible; Onlive has been doing it for some time already. This would be an interesting alternative for people that don't want to spend a lot of money up front on a gaming caliber box, but still be able to play "AAA" titles provided they aren't bothered by the I/O latency. This space would be a natural transition for Valve's gaming model, and create all sorts of new opportunities for charging models such as rental/pay by the hour without leaving your living room.

Note too that while Valve has stated they are working closely with both AMD and Nvidia, the first announced Steam Machine platform directly from Valve includes Nvidia hardware.  Keep in mind Nvidia already offers cloud gaming solutions and has the local streaming thing pretty much figured out (running on Linux no less). It would seem all the pieces may already be there. Microsoft and Sony are obviously in the race on this one as well.

The Performance Problem/Possibilities


Currently, the performance picture on Linux is somewhat spotty. There have been strides made in the last few months but the performance still generally trails its Windows counterpart. These performance differences are due almost entirely to drivers from hardware manufacturers. Of those, video card drivers have the most impact. For further reading on the current state of things, see: Nvidia, AMD, Intel. As for the future, there have already been announcements pledging better support.

Most games on the Windows platform utilize the API Direct3d, which acts as a unifying layer between the game and the video card. This allows for the same code to run on hardware from different vendors (in theory) without issue. This API is created by Microsoft, so it is unavailable on the Linux (SteamOS, Linux, MacOSX) platform. The standard API utilized by those platforms for 3d acceleration is and has been OpenGL for quite some time, which also runs on Windows. This would be the obvious choice for games going forward should they want to target both major PC platforms.

Oddly though, that isn't what a lot of developers would want to target anyhow. The more potentially lucrative pursuit would be to target consoles and the PC space, and anything coded to OpenGL and, to a lesser extent, Direct3d, would need work to be ported to consoles.

And now there's Mantle.

Mantle is the product of AMD's "next generation" console wins. The AMD Jaguar platform powers both the PS4 and Xbox One. To facilitate a unified API level and expose all the power possible on the platform, AMD created the "Mantle" API. Now AMD plans to bring this API to the PC. (Windows first, but it could work on any OS) It has already been adopted by the folks at Dice and they reported substantial performance gains. We'll have to wait for December to find out, but it will be very interesting to see if these performance promises pan out.

The potential here, however, is the possibility that this could further upset the entrenched standards in the industry. I don't think a hardware proprietary API is a good solution for the gaming community (remember Glide?), but if it catches on in any way the competition will be forced to respond in kind. AMD claims they'll open the platform eventually but I'd wager that will end up like the Nvidia "offer" to share PhysX. If the performance promises do pan out, one has to expect that Gabe (Valve CEO) would love to nudge its adoption if only to further undermine Direct3d.


Unintended Consequences


With all of the pressures introduced by the SteamOS introduction, Windows 8 lukewarm reception, the rise of the mobile market, and everything else mentioned above there will undoubtedly be changes across the board in an attempt to gain new or retain old customers. All these factors combined lead us to...

The Point.

So why do we care about this?


The pace of PC innovation hasn't kept up with other similar markets and it is now at its weakest point since its inception. Microsoft has done a poor job of keeping focus on the Windows based gaming space. Valve has introduced not only a hardware spec but a potential sans-Microsoft OS solution. Input, output, and consumption models are on the verge of being flipped on their head, and hardware manufacturers are moving to take more control of the API space. The probability of change to this platform stands higher than ever.

We don't know what will happen. For the first time in many years we don't know what will happen. The opportunities are many and the space is an incredibly exciting one to witness and participate in.

And this is the point. I'm intrigued by what this market could become for the first time in a long, long time. Perhaps Microsoft will re-invigorate the DirectX platform, or maybe SteamOS will finally bring the "year of the Linux desktop". Maybe the Oculus Rift combined with a Kinnect will launch the Xbox One to dominance, or maybe Nvidia releases a competing API that obsoletes everything prior.

Or maybe there's room for all of this. The market is growing and we're finally getting some competition in the field that is long overdue.

In times of rapid change the smallest actions can have large, lasting impacts. Even your product selection will impact the direction of video gaming in a small way, and any contributions you make to gaming communities even more-so. Now is our time to choose the direction, so go and contribute what you can. I know we'll all have fun along the way.

Wednesday, September 25, 2013

Overclocking Haswell Quick Fail Method

Overclocking is the process of increasing the clock speed on a system processor with the idea of getting higher performance. Given the time, this can be an interesting pursuit if you like pushing hardware performance to the highest level possible. With proper testing, this can be done once on a system and it will remain stable for 24/7/365 use for the life of the platform.

With a week off of work I decided to upgrade my aging i7-920 x58 platform to a Haswell based i7-4770k/z87. Overclocking this platform proved to have a steep learning curve, so I figured I'd share my general approach so others could save time getting started.

Overclocking Quick Fail Method

The general idea of the quick fail method (not sure if that's a name, but let's just pretend it is now) is to isolate each component of your build and then push it quickly to failure, then back it off and find the appropriate voltage/speed combo.


Overclocking is hard with the Paparazzi on your tail. 


Software/Hardware Needed


  • ADIA64 for Voltage/Temp monitoring
  • Prime95 v.27.9 for Stress Testing (Discussion Below); Alternatively you can use OCCT.
  • Intel K series processor (K series has an unlocked multiplier)
  • z87 chipset motherboard that supports overclocking; I used Asus but there are good boards from Gigabyte, MSI, and others as well. 

Getting Started/Clocking Approach


Make Sure you have applied your thermal interface material correctly and disabled the IGP in the BIOS(assuming you're not using it).  After your hardware is ready to go, we need to isolate components for testing. By doing so, we can more effectively ensure that each piece is stable before we combine them. We'll tweak/test in the following order:

  • Core
  • Uncore (Cache & Mem controller)
  • Memory
  • Combined

Those are also listed in order of performance. The Uncore, for example, does not need to be synced to the core clock and is not nearly as important from a performance perspective. 

Familiarize Yourself With Voltage and Temperature Norms


Before getting started, I recommend familiarizing yourself with the voltage/temperature behavior of you CPU at stock. To do so, we'll use the ADIA64 sensor display and the built in stress-test. (Note I don't recommend ADIA for stress testing other than here to gen heat/voltage) Launch ADIA64 and open the sensor display by expanding Computer->"Sensor". Take note of the CPU Core (vCore), CPU Cache (vCache, sometimes called vRing), VCCSA (or Vsa, System Agent), and CPU VRM (vRIN) settings, as we'll be targeting those for modification later. Launch the system stability test tool by selecting "Tools"->"System Stability Test" and click "Start". Watch the behavior of the voltages and take note of which are dynamic, and how they behave under load, and what they top out at. You'll need to know this for comparison purposes when we start adjusting.


Setting Voltages






Most z87 motherboards offer three or four options for overriding voltages: Offset, Adaptive, Manual and Auto. Let's take a moment to discuss each:

  • Auto: It's wise to avoid Auto unless you don't plan on spending much of any time testing. Most boards will apply far too much voltage to many aspects of the chip, creating additional heat which will waste energy and negatively effect your clocking potential. 
  • Manual will set voltage to a static number, which I'd shy away from as well for regular 24/7 use as it will counter alot of the benefit of the Haswell chip by not allowing for a dynamic reduction in voltage. Reducing the voltage reduces power consumption and heat mainly while idle. 
  • Offset: Now we're getting somewhere... Offset will apply a modifier to the target voltage on top of Intel's scaling algorithm. For example, if the vCore were set to .95 at a given point in time, and you specified a +.10 offset, then the resulting voltage would be 1.05. The only downside here is that the offset is always applied, but at least the voltage can scale downward. That leads us to:
  • Adaptive adds additional voltage on top of the offset voltage, but only while the turbo multiplier is active. This is a great new addition as it allows whatever is set here to be entirely rolled off when not in turbo mode. Unfortunately, you will still need to use offset for most of the voltage boost in most cases to keep things stable. A good rule of thumb to start with is putting the desired extra voltage 75% into the offset and 25% into the turbo offset. 
Note that many have found that adjusting the LLC ramp is not recommended when using Offset/Adaptive.

Stress Testing, Throttling, and the AVX Problem


Why Prime95?


Of all the tools I've tried, including ADIA64, Prime95 has provided me the most consistent and reproducible results. Prime seems to do a better job of bombing an unstable Haswell than a lot of other tools. For example, I was able to complete an ADIA64 stress test for two hours while the same setup locked up my system nearly immediately on Prime95SmallFFT.

More importantly, Prime allows for an effective customization of what portions of the CPU are being tested. SmallFFT will test the CPU and a bit of the cache(s) while Blend will effectively test the uncore and quite a bit of the system memory.


Haswell Thermal Throttling


The Haswell has an internal throttling mechanism (Thermal Control Circuit) to protect the chip that cannot be disabled. This throttling reduces the speed of the processor until the temperature drops below the temperature specified by Intel. For more information on the throttling point (at most 100 deg C) see table 26 in this document from Intel.

The AVX Issue


AVX are a series of extensions originally introduced in the SandyBridge platform that facilitate accelerating certain math functions. The problem with AVX is that when utilized by stress testing programs they're generally used in an "unrealistic" way (much more often than one would see) generating substantially more heat than the non-AVX equivalent processing. This can be enough heat to trigger thermal throttling of the Haswell chip which invalidates stress testing because it causes the chip to run at a slower speed.

We'll start testing with AVX enabled because we do want to ensure the cooling solution can handle as much heat as theoretically could be thrown out, but if it causes the chip to throttle we will disable it. That will be covered below.

Determine Max Core Speed


Lower Speed of Non-Core Components


To ensure we can definitively find the failure point of the core we need to take the uncore and memory out of the picture.
  1. Reduce the "Uncore" (sometimes called CPU Cache) multiplier to something substantially (5-8 depending on how aggressive) lower than your target overclock. For example if you are targeting 4.5Ghz, set the max Uncore multiplier to 38.  If there is an option for minimum Uncore multiplier, leave it at "Auto".
  2. Reduce the memory speed to something your RAM can easily handle, i.e. if you have DDR-2133, set it to DDR-1600 speed.  

Overclock the Core


Primary Voltages: vCore
Secondary Voltages: vRIN, vCache
Test: Prime95 SmallFFT





  1. Up the multiplier of the cores (if you have the option just sync all) to a reasonable starting point. On the i7-4770, for example, a good starting point would be 40 or so. 
  2. To ensure that your system isn't being throttled launch the ADIA64 stress testing tool but don't start it. The lower portion of the display will monitor and alert on throttling.
  3. Perform a 20 minute SmallFFT stress test using Prime95. To do so, launch the 64 bit Prime95 client and select the "SmallFFT" option. Make sure you monitor and take notes of the system voltages and temperatures during the test. If there is a problem with the overclock at this stage, it will generally manifest itself as a BSOD.  If your system throttles more than occasionally you need to disable AVX (see Appendix A) and re-start testing.
  4. If the system seems stable thus far feel free to up the multiplier and repeat. If the system crashed and you want to push further with the understanding that with more voltage and speed comes more power consumption and heat, go ahead and add voltage (see above for methods) and then repeat testing. Voltage to vCore should be added in steps of .01v to .02v or so. I wouldn't exceed 1.35v under load for 24/7 running. If vCore doesn't seem to have an impact here, you may want to bump CPU Input Voltage (vRIN) or the uncore voltage (vCache) just a bit... just don't go to high with those because they don't matter nearly as much as vCore. 
  5. After working your multiplier and voltage up to where you're willing to go, perform an extended stability test by saving the BIOS settings and then running the Prime95 SmallFFT test for at least 4 hours. If the extended test fails, bump voltage up or multiplier down and re-start testing. 

Overclock the UnCore


Primary Voltages: vCache
Secondary Voltages: vRIN, vCore, vSCCA
Tests: Prime95 SmallFFT, Prime95 Blend

Use the same methodology listed above to clock the Uncore. The following modifications exist: the main voltage to change is vCache rather than vCore. To ensure the uncore is stable you'll need to run Prime95 Blend after you finish the SmallFFT test. Don't be surprised if SmallFFT is stable and Blend requires more voltage and/or lowering the multiplier.

Clock Your Memory


Primary Voltages: vDIMM
Secondary Voltages: vSCCA, vRIN, vCache
Test: Prime95 Blend

Finally, clock your RAM using the correct voltage, speed, and timing adjustments. (Details outside the scope of this article) Use an extended Prime95 Blend test to finalize everything. Personally I'm not comfortable with anything less than 24 hrs for the final test, but opinions on that vary.


 Results/Final Thoughts


Obviously there is a ton here I'm not covering, but hopefully this information will save you time when starting your overclock. In some cases you may need to manipulate other voltages and voltage ramps, but those vary between motherboards so I won't touch on them.

Using this methodology I was able to get my water-cooled i4770k 24/7 stable @ 4.8Ghz/4/4Ghz Uncore without de-lidding. Not too shabby I should think. Finally I can get the performance I've been looking for when running Lotus 1-2-3.




Appendix A: (If Needed) Disable AVX

  1. Launch command prompt as administrator
  2. Execute "bcdedit /set xsavedisable 1"
  3. Reboot 
Remember to re-enable when done testing by using the same methodology, only executing bcdedit /set xsavedisable 0 rather than 1. 

Appendix B: General Usage Notes


  • If you want to take full advantage of Haswell's power management make sure you use a power profile for 24/7 use that includes reducing the core processor speed. (Stored in advanced options)
  • If using the IGP for media playback applications, note that the Haswell doesn't display the full 0-255 Quantization Range without major workarounds. 


Additional Reading

Sunday, August 18, 2013

Hyper-V Port Mirroring and Network Capture

Introduction

Hyper-V port mirroring, introduced in Windows Server 2012, allows you to easily monitor traffic on virtual machines without having to capture the traffic directly on that VM. It's exactly the same thing as standard port mirroring, only we're doing it using a virtual switch rather than a physical one. You can use this in your Hyper-V environment for all sorts of network troubleshooting. It only takes a couple of minutes to setup, so let's do it and then capture some packets!

Assumptions

  • Administrative Access to 2012 Hyper-V Host/Cluster
  • At least one vSwitch already configured (though we'll cover multiple)
  • At least two VMs configured; one to monitor and one to collect the data

Set Up Port Mirroring (GUI)


Configure Machine to be Mirrored

  1. Open the Hyper-V Manager
  2. Right click the machine that you would like to capture from and select "Settings".


  3. Expand the properties of the NIC (Network Interface Card) you would like to mirror by clicking the plus sign to its left and click "Advanced Features". 
  4. Under "Port mirroring"->"Mirroring Mode" click the drop down and select "Source". This sets this NIC as the source of mirroring on the Hyper-V switch it is connected to.


  5. Make note of the vSwitch the NIC is connected to (right below "Network Adapter") and click "OK".

Configure the Mirror Target


It makes setting up a network capture substantially easier if you add a dedicated NIC for each source machine. This NIC must be on the same virtual switch. A dedicated NIC allows unbinding all services/protocols in the guest OS, which will facilitate an entirely clean capture. More on that below...

  1. Shut down the virtual machine you intend on being the target of the network capture.
  2. After shutdown, right click that machine and select "Settings".
  3. Under "Add Hardware" in the right hand plane select "Network Adapter" and click "Add".
  4. The network adapter properties page for the new NIC will come up. Under "Virtual switch:" select the same switch that the source machine/NIC is connected to.
  5. Expand the properties of the new NIC and select "Advanced Features".
  6. Under "Port mirroring"->"Mirroring Mode" click the drop down and select "Destination". This sets this NIC as the source of mirroring on the Hyper-V switch it is connected to. Click "OK".


  7. Start the VM back up.

Configure the Mirror NIC in Capturing VM OS


Note: The instructions for this portion are somewhat generic because your guest OS on the capturing VM may differ from mine.

  1. After the capturing VM starts back up, log on via RDP or otherwise.
  2. Open your network connections and determine which NIC is the added one for mirroring. Rename it something like "{VSwitchName} Port Mirror" for easy identification.
  3. Open the properties of that NIC.
  4. Un-bind all protocols and services from that NIC and click "OK". By removing all bindings we'll be able to ensure a clean capture without interfering with the existing network connection. None of the standard protocols or services are used in the mirror process; Hyper-V takes care of everything for you. If you already have your network sniffing software installed, you may need to reboot the capture machine in order to see the NIC.


Install and Use Packet Capture Software


I'll be using Wireshark for Windows; if desired you could substitute something like Microsoft Network Monitor or Microsoft Message Analyzer on Windows; Wireshark or tcpdump on Linux. 

  1. Download and install Wireshark. The portable version works just as well if you prefer.
  2. Open Wireshark and click the "Interface List" button on the upper left hand corner.


  3. Select the dedicated capture NIC (which we renamed earlier), ensure it is the only selected, and click "Start".


  4. Enjoy all your packets.

Powershell Commands for NIC Setup/Mirroring

Caveats

  • You won't be able to decrypt encrypted packets unless you get the private key from the target server for decryption, which obviously may be a security issue given we're not on that machine.
  • Make sure you de-config (new word!) the port mirroring in HyperV when you're done as the packet replication continues even if you're not capturing.
  • After unbinding all services/protocols in Windows the adapter won't appear in the "Network and Sharing Center" anymore. You'll have to click "Change adapter settings" to get to the NIC.

More Reading



That's it, easy eh? Questions/Comments, leave 'em below!

Thursday, August 8, 2013

Starting Small: Set Up a Hadoop Compute Cluster Using Raspberry Pis

What is Hadoop?


Hadoop is a big data computing framework that generally refers to the main components: the core, HDFS, and MapReduce. There are several other projects under the umbrella as well. For more information, see this interview with Cloudera CSO Mike Olson.

What is a Raspberry Pi?


The Pi is a small, inexpensive ($39) ARM based computer. It is meant primarily as an educational tool.

Is Hadoop on the Pi practical?


Nope! Compute performance is horrendous.

Then why?


It's a great learning opportunity to work with Hadoop and multiple nodes. It's also cool to be able to put your "compute cluster" in a lunch box.
In reality though, this article is much more about setting up your first Hadoop compute cluster than it is about the Pi.

Could I use this guide to setup a non-Raspberry Pi based (Real) Hadoop Cluster?


Absolutely, please do. I'll make notes to that effect throughout.

I've been wanting to do this anyhow, let's get started!


Yeah, that's what I said.


Hardware Used


  • 3x Raspberry Pis
  • 1x 110v Power Splitter
  • 1x 4 port 10/100 Switch
  • 1x Powered 7 port USB hub
  • 3x 2' Cat 5 cables
  • 3x 2' USB "power" cables
  • 3x Class 10 16GB SDHC Cards

Total cost: about $170 bucks. I picked it all up at my local Microcenter, including the Pis!

Pi Hadoop Cluster; caution: may be slow.

If you would like a a high performance Hadoop cluster just pick one of these up on the way home:

This will cost at least $59.95. Maybe more. Probably more.


Raspberry Pi Preparation


We'll do the master node by itself first so that this guide can be used for a single or multi-node setup. Note: On a non-Pi installation skip to Networking Preparation (though you may want to update your OS manually).

Initial Config

  1. Download the "Raspbian wheezy" image from here and follow these directions to set up your first SD card. Soft-float isn't necessary.
  2. Hook up a monitor and keyboard and start up the device by connecting up the USB power.
  3. When the Raspberry Pi config tool launches, change the keyboard layout first. If you set passwords, etc. with the wrong KB layout you may have a hard time with that later.
  4. Change the timezone, then the locale. (Language, etc.)
  5. Change the default user password
  6. Configure the Pi to disable the X environment on boot. (boot_behavior->straight to desktop->no)
  7. Enable the SSH Daemon (Advanced-> A4 SSH) then exit the raspi-config tool
  8. Reboot the Pi. (sudo reboot)

Update The OS

Note: It is assumed from henceforth you are connecting to the Pi or your OS via SSH. You can continue on the direct terminal if you like until we get to multiple nodes.
  1. Log on to your Pi using the Raspberry user with the password you set earlier.
  2. To pull the newest sources, execute sudo apt-get update
  3. To pull upgrades for the OS and packages, execute sudo apt-get upgrade
  4. Reboot the pi. (sudo reboot)

Split the Memory/Overclock

  1. After logging back into the Pi, execute sudo raspi-config to bring up the Rasperry Pi config tool.
  2. Select Option 8, "Advanced Options"
  3. Select Option A3, "Memory Split"
  4. When prompted for how much memory the GPU should have, enter the minimum, "16" and hit enter.
  5. (If you would like to overclock) Select Option 7, "Overclock", hit "OK" to the warning, and select what speed you would like.
  6. On the main menu, select "finish" and hit "Yes" to reboot.

Networking Preparation

Each Hadoop node must have a unique name and static IP.  I'll be following the Debian instructions since the Rasbian build is based on Debian. If you're running another distro follow the instructions for that. (Redhat, CentOS, Ubuntu)

First we need to change the machine hostname by editing /etc/hostname
sudo nano /etc/hostname
Note: I'll be using nano throughout the article; I'm sure some of you will be substituting vi instead. :)
Change the hostname to what you would like and save the file; I used "Node 1" throughout this document. If you're setting up the second node make sure to use a different hostname.

Now to change to a static IP by editing /etc/network/interfaces
sudo nano /etc/network/interfaces
Change the iface eth0 to a static IP by replacing "auto eth0" and/or "iface eth0 inet dhcp" with:
iface eth0 inet static
        address 192.168.1.40
        netmask 255.255.255.0
        gateway 192.168.1.1
Note: Substitute the correct address (IP), netmask, and gateway for your environment.

Make sure you can resolve DNS queries correctly by editing your /etc/resolv.conf:
sudo nano /etc/resolv.conf
Change the content to match below, substituting the correct information for your environment:
domain company.com
search company.com
nameserver 192.168.1.20
nameserver 192.168.1.30
domain=domain suffix for this machine
search=appends when FQDN not specified
nameserver=list DNS servers in order of precedence

We could restart networking etc, but for the sake of simplicity we'll bounce the box:
sudo reboot

Install Java

If you're going for performance you can install the SunJDK, but getting that to work requires a bit of extra effort. Since we're on the Pi and performance isn't our goal, I'll be using OpenJDK which installs easily. If you're installing on a "real" machine/VM, you may want to diverge a bit here and go for the real deal (Probably Java ver 6).
After connecting to the new IP via SSH again, execute:
sudo apt-get install openjdk-7-jdk
Ensure that 7 is the version you want. If not substitute the package accordingly.




Create and Config Hadoop User and Groups

Create the hadoop group:
sudo addgroup hadoop
Create hadoop user "hduser" and place it in the hadoop group; make sure you remember the password!
sudo adduser --ingroup hadoop hduser
Give hduser the ability to sudo:
sudo adduser hduser sudo

Now logout:
logout
Log back in as the newly created hduser. From hence forth everything will be executed as that user.

Setup SSH Certs

SSH is used between Hadoop nodes and services (even on the same node) to coordinate work. These SSH sessions will be run as the Hadoop user, "hduser" that we created earlier. We need to generate a keypair to use for EACH node. Let's do this one first:

Create the certs for use with ssh:
ssh-keygen -t rsa -P ""
When prompted, save the key to the default directory (will be /home/USERNAME/.ssh/id_rsa). If done correctly you will be shown the fingerprint and randomart image.


Copy the public key into the user's authorized keys store (~ represents the home directory of the current user, >> appends to that file to preserve any keys already there) 
cat ~/.ssh/id_rsa.pub >> ~/.ssh/authorized_keys
SSH to localhost and add to the list of known hosts.
ssh localhost
When prompted, type "yes" to add to the list of known hosts.

Download/Install Hadoop!

Note: As of this writing, the newest version of the 1.x branch (which we're covering) is 1.2. (Ahem, 1.2.1, they're quick) You should check to see what the newest stable release is and change the download link below accordingly. Versions and mirrors can be found here.

Download to home dir:
cd ~
wget http://mirror.catn.com/pub/apache/hadoop/core/hadoop-1.2.0/hadoop-1.2.0.tar.gz
Unzip to the /usr/local dir:
sudo tar vxzf hadoop-1.2.0.tar.gz -C /usr/local
Change the dir name of the unzipped hadoop dir. Note that if your version # is different you'll need to adjust the command
cd /usr/local
sudo mv hadoop-1.2.0/ hadoop
Set the hadoop dir owner to be the hadoop user (hduser):
sudo chown -R hduser:hadoop hadoop

Configure the User Environment

Now we'll config the environment for the user (hduser) to run Hadoop. This assumes you're logged on as that user.
Change to the home dir and open the .bashrc file for editing:
cd ~
nano .bashrc
Add the following to the .bashrc file. It doesn't matter where: (note if you're using Sun Java on a real box you need to put in the right dir for JAVA_HOME)
export JAVA_HOME=/usr/lib/jvm/java-7-openjdk-armhf
export HADOOP_INSTALL=/usr/local/hadoop
export PATH=$PATH:$HADOOP_INSTALL/bin
Bounce the box (you could actually just log off/on, but what the hell eh?):
sudo reboot

Configure Hadoop!

(Finally!) It's time to setup Hadoop on the machine. With these steps we'll set up a single node and we'll discuss multiple node configuration afterward.
First let's ensure your Hadoop install is in place correctly. After logging in as hduser, execute:
hadoop version
you should see something like: (differs depending on version)
Hadoop 1.2.0
Subversion https://svn.apache.org/repos/asf/hadoop/common/branches/branch-1.2 -r 1479473
Compiled by hortonfo on Mon May  6 06:59:37 UTC 2013
From source with checksum 2e0dac51ede113c1f2ca8e7d82fb3405
This command was run using /usr/local/hadoop/hadoop-core-1.2.0.jar
Good. Now let's configure. First up we'll set up the environment file that defines some overall runtime parameters:
nano /usr/local/hadoop/conf/hadoop-env.sh
Add the following lines to set the Java Home parameter (specifies where Java is) and how much memory the Hadoop processes are allowed to use. There should be sample lines in your existing hadoop-env.sh that you can just un-comment and modify.
Note: If you're using a "real" machine rather than a Pi, change the heapsize to a number appropriate to your machine (physical mem - (OS + caching overhead)) and ensure you have the right directory for JAVA_HOME because it will likely be different. If you have a Raspberry Pi version "A" rather than "B", you'll heapsize will need to be much lower since the RAM is halved.
export JAVA_HOME=/usr/lib/jvm/java-7-openjdk-armhf
export HADOOP_HEAPSIZE=272
Now let's configure core-site.xml. This file defines critical operational parameters for a Hadoop site.
Note: You'll notice we're using "localhost" in the configuration. That will change when we go to multi-node, but we'll leave it as localhost for a demonstration of an exportable local only configuration.
nano /usr/local/hadoop/conf/core-site.xml
Add the following lines. If there is already information in the file make sure you respect XML format rules. The "description" field is not required, but I've included it to help this tutorial make sense.
<configuration>
  <property>
    <name>hadoop.tmp.dir</name>
    <value>/fs/hadoop/tmp</value>    <description>Sets the operating directory for Hadoop data.
    </description>
  </property>
  <property>
    <name>fs.default.name</name>
    <value>hdfs://localhost:54310</value>    <description>The name of the default file system.  A URI whose
    scheme and authority determine the FileSystem implementation.
    The URI's scheme determines the config property (fs.SCHEME.impl) naming
    the FileSystem implementation class.  The URI's authority is used to
    determine the host, port, etc. for a filesystem.  
    </description>
  </property>
</configuration>
It's time to configure the mapred-site.xml. This file determines where the mapreduce job tracker(s) run. Again, we're using localhost for now and we'll change it when we go to multi-node in a bit.
nano /usr/local/hadoop/conf/mapred-site.xml
Add the following lines:
<configuration>
  <property>
    <name>mapred.job.tracker</name>
    <value>localhost:54311</value>
    <description>The host and port that the MapReduce job tracker runs
    at.  If "local", then jobs are run in-process as a single map
    and reduce task.
    </description>
  </property>
</configuration>
Now on to hdfs-site.xml. This file configures the HDFS parameters.
nano /usr/local/hadoop/conf/hdfs-site.xml
Add the following lines:
Note: The dfs.replication value sets how many copies of any given file should exist in the configured HDFS site. We'll do 1 for now since there is only one node, but again when we configure multiple nodes we'll change this.
<configuration>
  <property>
    <name>dfs.replication</name>
    <value>1</value>
    <description>Default block replication.
    The actual number of replications can be specified when the file is created.
    The default is used if replication is not specified in create time.
    </description>
  </property>
</configuration>
Create the working directory and set permissions
sudo mkdir -p /fs/hadoop/tmp
sudo chown hduser:hadoop /fs/hadoop/tmp
sudo chmod 750 /fs/hadoop/tmp/
Format the working directory
/usr/local/hadoop/bin/hadoop namenode -format

Start Hadoop and Run Your First Job

Let's roll, this will be fun.
Start it up:
cd /usr/local/hadoop
bin/start-all.sh
Now let's make sure it's running successfully. First we'll use the jps command (Java PS) to determine what Java processes are running:
jps
You should see the following processes: (ignore the number in front, that's a process ID)
4863 Jps
4003 SecondaryNameNode
4192 TaskTracker
3893 DataNode
3787 NameNode
4079 JobTracker
If there are any missing processes, you'll need to review logs. By default, logs are located in:
/usr/local/hadoop/logs
There are two types of log files: .out and .log. .out files detail process information while .log files are the logging output from the process. Generally .log files are used for troubleshooting. The logfile name standard is:
hadoop-(username)-(processtype)-(machinename).log
i.e.
hadoop-hduser-jobtracker-node1.log
In this, like many other articles, we'll use the included wordcount example. For the wordcount example you will need plain text large enough to be interesting. For this purpose, you can download plain text books from Project Gutenberg.

After downloading 1 or more books (make sure you selected "Plain Text UTF-8" format) copy them to the local filesystem on your Hadoop node using SSH (via SCP or similar). In my example I've copied the books to /tmp/books.

Now let's copy the books onto the HDFS filesystem by using the Hadoop dfs command.
cd /usr/local/hadoop
bin/hadoop dfs -copyFromLocal /tmp/books /fs/hduser/books
After copying in the book(s), execute the wordcount example:
bin/hadoop jar hadoop*examples*.jar wordcount /fs/hduser/books /fs/hduser/books-output
You'll see the job run; it may take awhile... remember these Pis don't perform all that well. Upon completion you should see something like this:
13/06/17 15:46:54 INFO mapred.JobClient: Job complete: job_201306170244_0001
13/06/17 15:46:55 INFO mapred.JobClient: Counters: 29
13/06/17 15:46:55 INFO mapred.JobClient:   Job Counters
13/06/17 15:46:55 INFO mapred.JobClient:     Launched reduce tasks=1
13/06/17 15:46:55 INFO mapred.JobClient:     SLOTS_MILLIS_MAPS=566320
13/06/17 15:46:55 INFO mapred.JobClient:     Total time spent by all reduces waiting after reserving slots (ms)=0
13/06/17 15:46:55 INFO mapred.JobClient:     Total time spent by all maps waiting after reserving slots (ms)=0
13/06/17 15:46:55 INFO mapred.JobClient:     Launched map tasks=1
13/06/17 15:46:55 INFO mapred.JobClient:     Data-local map tasks=1
13/06/17 15:46:55 INFO mapred.JobClient:     SLOTS_MILLIS_REDUCES=112626
13/06/17 15:46:55 INFO mapred.JobClient:   File Output Format Counters
13/06/17 15:46:55 INFO mapred.JobClient:     Bytes Written=229278
...
The results from the run, if you're curious, should be in the last part of the command you executed earlier. (/fs/hduser/books-output in our case) Update: to check the output use the hadoop dfs command a la:
hduser@rasdoop1 /usr/local/hadoop $ bin/hadoop dfs -ls /fs/hduser/books-output
Found 3 items
-rw-r--r--   3 hduser supergroup          0 2015-05-05 05:00 /fs/hduser/books-output/_SUCCESS
drwxr-xr-x   - hduser supergroup          0 2015-05-05 04:35 /fs/hduser/books-output/_logs
-rw-r--r--   3 hduser supergroup     926451 2015-05-05 04:58 /fs/hduser/books-output/part-r-00000

hduser@rasdoop1 /usr/local/hadoop $bin/hadoop dfs -cat /fs/hduser/books-output/part-r-00000
...
Congrats! You just ran your first Hadoop job, welcome to the world of (tiny)big data!

Multiple Nodes


Hadoop has been successfully run on thousands of nodes; this is where we derive the power of the platform. The first step to getting to many nodes is getting to 2, so let's do that.

Set up the Secondary Pi

I'll cover cloning a node for the third Pi, so for this one we'll assume you've set up another Pi the same as the first using the instructions above. After that, we just need make appropriate changes to the configuration files. Important Note: I'll be referring to the nodes as Node 1 and Node 2 and making the assumption that you have named them that. Feel free to use different names, just make sure you substitute them in below. Name resolution will be an issue; we'll touch on that below.

Node 1 will run:
  • NameNode
  • Secondary NameNode (In a large production cluster this would be somewhere else)
  • DataNode
  • JobTracker (In a large production cluster this would be somewhere else)
  • TaskTracker
Node 2 will run:
  • DataNode
  • TaskTracker
Normally this is where I planned on writing up an overview of the different pieces of Hadoop, but upon researching I found an article that exceeded what I planned to write it by such a magnitude I figured why bother, I'll just link it here. Excellent post by Brad Hedlund, be sure to check it out.
On Node 1, edit masters
nano /usr/local/hadoop/conf/masters
remove "localhost" and add the first node FQDN and save the file:
Node1.domain.ext
Note that conf/masters does NOT determine which node holds "master" roles, i.e. NameNode, JobTracker, etc. It only defines which node will attempt to contact the nodes in slaves (below) to initiate start-up.
On Node 1, edit slaves
nano /usr/local/hadoop/conf/slaves
remove "localhost" add the internal FQDN of all nodes in the cluster and save the file:
Node1.domain.ext
Node2.domain.ext
On ALL (both) cluster members edit core-site.xml
nano /usr/local/hadoop/conf/core-site.xml
and change the fs.default.name value to reflect the location of the NameNode:
<property>
 <name>fs.default.name</name>
 <value>hdfs://node1:54310</value>
</property>
On ALL (both) cluster members edit mapred-site.xml
nano /usr/local/hadoop/conf/mapred-site.xml
and change the mapred.job.tracker value to reflect the location of the JobTracker:
<property>
 <name>mapred.job.tracker</name>
 <value>node1:54311</value>
</property>
On ALL (both) cluster members edit hdfs-site.xml
nano /usr/local/hadoop/conf/hdfs-site.xml
and change the dfs.replication value to reflect the location of the JobTracker:
<property>
 <name>dfs.replication</name>
 <value>2</value>
</property>
The dfs.replication value determines how many copies of any given piece of data will exist in the HDFS site. By default this is set to 3 which is optimal for many small->medium clusters. In this case we set it to the number of nodes we'll have right now, 2.

Now we need to ensure the master node can talk to the slave node(s) by adding the master key to the authorized keys file on each node (just 1 for now)
On Node 1, cat your public key and copy it to the clipboard or some other medium for transfer to the slave node
cat ~/.ssh/id_rsa.pub
On Node 2
nano ~/.ssh/authorized_keys
and paste in the key from the master and save the file to allow it to take commands.

Name Lookup

The Hadoop nodes should be able to resolve the names of the nodes in the cluster. To accomplish this we have two options: the right way and the quick way.

The right way: DNS

The best way to enable name lookup works correctly is by adding the nodes and their IPs to your internal DNS. Not only does this facilitate standard Hadoop operation, but it makes it possible to browse around the Hadoop status web pages (more on that below) from a non-node member, which I suspect you'll want to do. Those sites are automatically created by Hadoop all the hyperlinks use the node DNS names, so you need to be able to resolve those names from a "client" machine. This can scale too; with the right tool set it's very easy to automate DNS updates when spinning up a node.

To do this option, add your Hadoop hostnames to the DNS zone they reside in.

The quick way: HOSTS

If you don't have access to update your internal DNS zone, you'll need to use the hosts files on the nodes. This will allow the nodes to talk to each other via name. There are scenarios where you may want to automate the updating of node based HOSTS files for performance or fault tolerance reasons as well, but I won't go into that here.

To use this option do the following:

On ALL (both) cluster members edit /etc/hosts
nano /etc/hosts
and add each of your nodes with the appropriate IP addresses (mine are only examples)
192.168.1.40    node1
192.168.1.41    node2
Now that the configuration is done we need to format the HDFS filesystem. Note: THIS IS DESTRUCTIVE. Anything currently on your single node Hadoop system will be deleted.
On Node 1
bin/hadoop namenode -format
When that is done, let's start the cluster:
cd /usr/local/hadoop
bin/start-all.sh
Just as before, use the jps command to check processes, but this time on both nodes.
On Node 1
jps
You should see the following processes: (ignore the number in front, that's a process ID)
4863 Jps
4003 SecondaryNameNode
4192 TaskTracker
3893 DataNode
3787 NameNode
4079 JobTracker
On Node 2
jps
You should see the following processes: (ignore the number in front, that's a process ID)
6365 TaskTracker
7248 Jps
6279 DataNode
If there are problems review the logs. (for detailed location see reference above) Note that by default each node contains its own logs.

Now let's run the wordcount example in a cluster. To facilitate this you'll need enough data to split, so when downloading UTF-8 books from Project Gutenberg get at least 6 very long books (that should do with default settings).

Copy them to the local filesystem on your master Hadoop node using SSH (via SCP or similar). In my example I've copied the books to /tmp/books.

Now let's copy the books onto the HDFS filesystem by using the Hadoop dfs command. This will automatically replicate data to all nodes.
cd /usr/local/hadoop
bin/hadoop dfs -copyFromLocal /tmp/books /fs/hduser/books
After copying in the book(s) wait a couple minutes for all blocks to replicate (more on monitoring below) then execute the wordcount example:
bin/hadoop jar hadoop*examples*.jar wordcount /fs/hduser/books /fs/hduser/books-output

Monitoring HDFS/Jobs


Hadoop includes a great set of management web sites that will allow you to monitor jobs, check logfiles, browse the filesystem, etc. Let's take a moment to examine the three sites.

JobTracker
By default, the job tracker site can be found at http://(jobtrackernodename.domain.ext):50030/jobtracker.jsp ; i.e. http://node1.company.com:50030/jobtracker.jsp



On the Job Tracker you can see information regarding the cluster, running map and reduce tasks, node status, running and completed jobs, and you can also drill into specific node and task status.

DFSHealth
By default, the DFSHealth site can be found at http://(NameNodename.domain.ext):50070/dfshealth.jsp ; i.e. http://node1.company.com:50070/dfshealth.jsp


The DFSHealth site allows you to browse the HDFS system, view nodes, look at space consumption, and check NameNode logs.

TaskTracker
By default, the task tracker site can be found at http://(nodename.domain.ext):50060/tasktracker.jsp ; i.e. http://node2.company.com:50060/tasktracker.jsp



The Task Tracker site runs on each node that can run a task to provide more detailed information about what task is running on that node at the time.


Cloning a Node and Adding it to the Cluster


Now we'll explore how to get our third node into the cluster the quickest way possible: cloning the drive. This will cover, at a high level, the SD card cloning process for the Pi. If you're using a real machine replace the drive cloning steps with your favorite (*cough*Clonzilla*cough*) cloning software, or copy the virtual drive if you are using VMs.

You should always clone a slave node unless you intend on making a new cluster. To this end, do the following:

  1. Shut down Node 2 (sudo init 0) and remove the SD card.
  2. Insert the SD card into a PC to so we can clone it; I'm using a Windows 7 machine so if you're using Linux we'll diverge here. (I'll meet you at the pass!)
  3. Capture an image to your hdd using your favorite SDcard imaging software. I use either Roadkill's Disk Image or dd for windows
  4. Remove the Node 2 SD card and place it back into Node 2. Do not power up yet.
  5. Insert a new SD card of the same size to your imaging PC.
  6. Delete any partitions on the card to ensure error-free imaging. I use Diskpart , Select Disk x , Clean , exit where "x" is the disk number. Use list disk to find it... don't screw up or you'll wipe out your HDD!
  7. Image the new SD card with the image from Node 2 and remove it from the PC when complete.


  8. Insert the newly cloned SD card to Node 3. Power up that node and leave Node 2 off for now or the IPs will conflict.

Log into the new node (same IP as the cloned node) via ssh with the hduser (Hadoop user) account and perform the following tasks:

Change the hostname:
sudo nano /etc/hostname
For this example change "Node2" to "Node3" and save the file.

Change the IP address:
sudo nano /etc/network/interfaces
For this example change "192.168.1.41" to "192.168.1.42" and save the file. Switch to your own IP address range if necessary.

If you're using HOSTS for name resolution, add Node 3 with the proper IP address to the hosts file
sudo nano /etc/hosts
Repeat this on the master node and for the sake of completeness the other node(s) should probably be updated as well. In a large scale deployment this would all be automated.
If you are using DNS, make sure to add the Node3 entry to your zone.

Reboot Node 3 to enact the IP and hostname changes
sudo reboot
Power up Node 2 again at this time if you had it powered down. Log back into Node 3 (now with the new IP!) with the hduser account via SSH.

Generate a new key:
ssh-keygen -t rsa -P ""

Copy the key to its own trusted store:
cat ~/.ssh/id_rsa.pub >> ~/.ssh/authorized_keys

If you aren't already, log into Node1 (Master) as hduser via SSH and do the following:
Add Node 3 to the slaves list
nano /usr/local/hadoop/conf/slaves

SSH to Node 3 to add it to known hosts
ssh rasdoop3
When prompted re: continue connecting type yes

An ALL nodes edit hdfs-site.xml
nano /usr/local/hadoop/conf/hdfs-site.xml
Change it to replicate to 3 nodes. Note if you add more nodes to the cluster after this you would most likely NOT increase this above 3. It's not 1:1 that we're going for, it's the number of replicas we want in the entire cluster. Once you're running a cluster with > 3 nodes you'll generally know if you want dfs.replication set to more than 3.
<configuration>
  <property>
    <name>dfs.replication</name>
    <value>3</value>
    <description>Default block replication.
    The actual number of replications can be specified when the file is created.
    The default is used if replication is not specified in create time.
    </description>
  </property>
</configuration>

Now we need to format the filesystem. It is possible, but slightly more complicated, to clone/add a node without re-formatting the whole HDFS cluster, but that's a topic for another blog post. Since we have virtually no data here, we'll address all concerns by wiping clean:
On Node 1:
cd /usr/local/hadoop
bin/hadoop namenode -format

Now we should be able to start up all three nodes. Let's give it a shot!
On Node 1:
bin/start-all.sh

Note: If you have issues starting up the Datanodes you may need to delete the sub-directories under the directory listed in core-site.xml as "hadoop.tmp.dir" and then re-format again (see above). 

As for startup troubleshooting and running the wordcount example, copy out the instructions above after the initial cluster setup. There should be nothing different from running with two nodes save the fact that you may need more books to get it working on all 3 nodes.


Postmortem/Additional Reading

You did it! You've now got the experience of setting up a Hadoop cluster under your belt. There's a ton of directions to go from here; this relatively new tech is changing the way a lot of companies view data and business intelligence.

There is so much great community content out there; here's just a small list of the references I used and some additional reading.

References:

Raspberry Pi forums: Hadoop + HDFS + MR on Pi cluster - works!
Michael G. Noll: Running Hadoop on Ubuntu Linux Single-Node Cluster
Michael G. Noll: Running Hadoop on Ubuntu Linux Multi-Note Cluster
Hadoop Wiki: Getting Started with Hadoop
Hadoop Wiki: Wordcount Example
University of Glasgow's Raspberry Pi Hadoop Project
Hadoop Wiki: Cluster Setup
Hadoop Wiki: Java Versions
All Things Hadoop: Tips, Tricks and Pointers When Setting Up..
Hadoop Wiki: FAQ
Code/Google Hadoop Toolkit: Hadoop Performance Monitoring
Jeremy Morgan: How to Overclock the Raspberry Pi