- はてなMobileGatewayを使用
- ポップアップにてgmail送信ウィンドウが開く
ブックマークレットの作り方
-
↓ここに自分の携帯メールアドレスを入れて
(ブラウザ上で処理されるため、入力内容は当サービスに送信されません。) -
-
と、↓ここに [携帯で読む] というリンクが表示されるので
-
できたリンクを右クリックして
ブックマークなりお気に入りなりに追加する。
(「このリンクは安全でない可能性が」などと出ても気にしない。) なお、動作確認FireFox3でのみ行っています。
↓ここに自分の携帯メールアドレスを入れて
(ブラウザ上で処理されるため、入力内容は当サービスに送信されません。)
と、↓ここに [携帯で読む] というリンクが表示されるので
できたリンクを右クリックして
ブックマークなりお気に入りなりに追加する。
(「このリンクは安全でない可能性が」などと出ても気にしない。) なお、動作確認FireFox3でのみ行っています。
Today, Opera revealed the newest version of their web browser, Opera 9.6. As always, the latest update includes speed and performance increases, but the update delivers several new features, too. The one new feature that we were really excited to try out is how Opera 9.6 deals with RSS feeds. In this latest version of the browser, you can preview your feeds in an attractive magazine-style layout. But what we really wanted to know is could read your feeds like this once subscribed?
In Opera 9.6, a new feed preview feature has been introduced that turns any RSS feed into a magazine-style page where the articles in an feed appear as columns. (See image below). With the feeds laid out in this manner, suddenly RSS reading becomes accessible, understandable, and far less geeky than its acronym implies. Although heavy RSS users and techie folks will probably continue to use an RSS reader like Google Reader, a magazine-style layout is a great option for a light reader or someone new to RSS.
Previewing RWW's Feed
In a way, Opera's new magazine-style feature reminds us very much of how the Firefox extension, Feedly, operates. With Feedly installed, you can view your Google Reader feeds in an easy-to-read format while still being able to hop into your different folders. Of course, Feedly does so much more than just change the layout of Google Reader, but that's a whole other topic.
The difference between Feedly's magazine interface and what Opera does is that Opera only displays feeds in this manner when you preview them while deciding whether or not to subscribe. That's disappointing. We were hoping that Opera would include this as a new option under the "Display -> View" settings in Opera's built-in RSS reader, too. Unfortunately, those view settings have remained the same. Feed reading there is still an inbox-like experience, with feeds titles in one window and the articles in a second window. This familiar Outlook format works for some people, we're sure, but to have the magazine-style option here as well would have been a nice treat.
In addition to the preview feeds feature, Opera 9.6 also adds other updates, including the following:
In times of high stress, many in the financial world seek solace in watery metaphors. We hear of vast irresistible forces converging in "perfect storms" and unforeseeable events contributing to "100-year floods."
How could we have expected, let alone prevented, this?
Count on Warren E. Buffett to cut to the truth. Years ago, referring to reckless corporate debt, Buffett noted (or so the story goes): "You never know who is swimming naked until the tide goes out."
The tide's moving, and we're starting to get the full, not-so-pretty view. Along with the bare swimmers emerging from the soggy murk, we're being reminded of some of the dumb ideas and reckless choices that helped deliver us to our current debacle. As stunning as the scene seems, we've actually had plenty of experience with this sort of thing. But like some stubborn residents of hurricane zones, we swiftly choose to forget the last tempest and reassure ourselves that things will be different from now on. Why don't we learn the obvious lesson to the contrary? Answers: the timeless power of hubris during periods when profits seem easy, and a set of foolish financial notions that have become prevalent over the past three decades.
One of those beliefs is the indiscriminate antiregulatory ideology one hears preached on Wall Street with tent-revival fervor. What makes this thinking so perplexing is that many of the free-market true believers also assume the federal government will save them if they flop. Consider the extraordinary taxpayer-backed rescues of insurance titan American International Group (AIG), housing financiers Fannie Mae (FNM) and Freddie Mac (FRE), and, before those, the Treasury-guided merger of Bear Stearns into JPMorgan Chase (JPM). It brings to mind the homeowner who rants about getting Washington off his back but wants federally guaranteed flood insurance no matter how close to the Gulf Coast he builds his house.
Other by-now-familiar attitudes have helped put us in the drink: In good times, there's no such thing as too much leverage. (Remember Michael Milken?) Derivatives don't require oversight, even though almost no one understands them. (How now, Long-Term Capital Management?) And, don't worry, the quantitative geniuses have devised models to eliminate extreme risk. (Enron, anyone?)
"Now, again, the banks and the Bush Administration and [Treasury Secretary Henry] Paulson and [Federal Reserve Chairman Ben] Bernanke would like you to think these crises are like floods or hurricanes," says Michael Greenberger, a senior official at the Commodity Futures Trading Commission (CFTC) during the Clinton Administration. An advocate of more aggressive regulation of investment banks, he was shot down in the late 1990s by Democratic colleagues, not just GOP foes. Most financial calamities aren't like natural forces beyond control, Greenberger says. "These are predictable events." Predictable events, of course, are more likely to be prevented with sound rules and stiff enforcement.
Different Animals
Alfred E. Kahn offers the long view—a very long view. As the Carter Administration's aviation czar, he unshackled airline routes and fares in the late 1970s, reshaping that industry (for better and worse) and helping spur a lengthy era of economic deregulation. Still sharp at 91, the retired Cornell University economist and part-time consultant recalls that almost as soon as the free-market spirits were set loose, a furious stampede ensued. Lenders, for one, demanded lots more freedom. But they "were a different kind of animal" from airlines and trucking firms, which the Carterites also deregulated, Kahn says. "They were animals that had a direct effect on the macroeconomy. That is very different from the regulation of industries that provided goods and services.…I never supported any type of deregulation of banking."
During the Reagan years, Kahn's cautious industry-by-industry analysis was replaced by the all-encompassing antiregulatory ideology of the University of Chicago. One result: the liberation of an armada of savings and loan pirates, abetted by congressional Democrats as well as Republicans, many of them drunk on S&L campaign largesse. (Wall Street lobbyists with open wallets have since perfected the practice of neutralizing Congress on a bipartisan basis.) Hundreds of thrifts ultimately collapsed in the late 1980s and 1990s amid greedy and, in some cases, fraudulent real estate deals.
As early as 2000, William J. Brennan, a prominent consumer attorney who has represented mortgage borrowers since the S&L catastrophe, warned in testimony before the House Financial Services Committee that real estate finance would return in new guises to haunt us. Few listened. Behind every burst of ill-advised lending lurk financial innovators creating new mechanisms to entice ever-more-sketchy borrowers, says Brennan, the director of Atlanta Legal Aid Society's Home Defense Program. In the 1980s, Michael Milken and his comrades at the now-defunct Drexel Burnham Lambert investment bank exacerbated the S&L fiasco by hawking their thrift clients' high-risk junk bonds. More recently the likes of soon-to-be-defunct Lehman Brothers and Bear Stearns engineered the securitization of mortgages, encouraging home lenders to spew wildly unwise loans. "Lending without regard [for] the ability to pay back started with the S&L scandal," says Brennan. In the 1980s the borrowers were reckless shopping-mall developers; in the recent boom, unsophisticated and sometimes cavalier homeowners.
Wall Street transformed dicey subprime mortgages into the toxic securities that have required hundreds of billions in writedowns and that drove once-mighty Merrill Lynch (MER) to sell itself to Bank of America (BAC). One of the most striking aspects of the current turbulence is the degree to which banks invested in the noxious fare themselves, notes Emanuel Derman, who heads risk management at Prisma Capital Partners, a hedge fund in Jersey City, N.J. "These guys ate their own cooking; they didn't just pass it on to clients."
The outsize appetite on Wall Street for hazardous mortgage-backed securities and even more obscure derivatives has had a lot to do with the people in the kitchen failing to understand fully what was in their recipes. All of this is painfully familiar to anyone who paid attention to past adventures with wizards who claimed their esoteric models had magically eliminated risk and uncertainty. Hedge fund Long-Term Capital Management (LTCM) couldn't imagine Russia defaulting on its debt, much as Lehman apparently couldn't conceive of housing prices across the country deteriorating simultaneously, followed by a paralyzing credit crunch.
For four years in the mid-1990s, LTCM boasted extraordinary profits based on supposedly flawless computer formulas devised by a team that included two Nobel laureates. But in the summer of 1998, Russian credit disintegrated, one of several concurrent global shocks that the LTCM crew had failed to factor into their algorithms. After losing more than $4 billion in a few months—in retrospect, the amount seems almost quaint—the hedge fund received a federally organized rescue, although it later shut down altogether.
Financial "rocket scientists," says Henry T. Hu, a corporate law professor at the University of Texas in Austin, have a knack for neglecting low-probability, catastrophic events. The smartest guys in the room at Enron similarly assumed away risks they didn't want to confront. "These models…work in normal circumstances but not during times of market stress, when it really matters," Hu says. "It is almost like a safety belt that only fails in a serious car crash."
One of the things that dismayed outsiders about LTCM after it came apart was the size and complexity of its derivatives portfolio. Some in the Clinton Administration pushed for more oversight of the unregulated, privately traded instruments whose value derives from price shifts in currencies, securities, or other assets. Then-Fed Chairman Alan Greenspan, allied with Robert E. Rubin, Clinton's Treasury Secretary (and now a director and senior counselor at Citigroup (C), opposed tougher policing of derivatives. Banks could watch over each other more effectively than regulators could, Greenspan argued. This turned out to be shortsighted.
In an interview, Greenspan doesn't back down, even after all we've seen lately. "The majority of lawyers, in my experience, seek to regulate—that is, to contain certain activities with little weight given to the lost benefits of such activities," he says. "The question is: What do you lose? In this case, a very valuable instrument [credit default swaps, the derivatives at the core of the current mess] for the diminution of systemic risk. You can stop the system dead and eliminate speculative losses. But you will also get significantly reduced economic activity and ultimately lower standards of living."
Greenspan adds: "I've been extraordinarily distressed by how badly the most sophisticated people in the business handled risk management. But the question is: If, protecting their own resources, they can't do it, who's going to do it better?" (Well, maybe regulators who don't have big bonuses at stake would be less likely to get carried away by the euphoria.)
Rubin says separately that he didn't oppose the general idea of scrutinizing derivatives, but instead argued against particular proposals in the late '90s to expand CFTC authority. "I have always been concerned about derivatives," he says.
Michael Greenberger served as the CFTC's director of trading and markets at the time. A proponent of tougher oversight, he recalls the Greenspan-Rubin resistance as being fierce and across-the-board. "If we had prevailed, the [subprime-securitization] party would never have gotten started; the wildness wouldn't have happened," he says. "There would have been auditing requirements, capital requirements, transparency. No more operating in the shadows. Bear Stearns, Lehman, Enron, and AIG would be thriving, and spending every waking hour complaining about regulatory restraints imposed upon them." Now a law professor at the University of Maryland, Greenberger adds: "In a booming economy, people couldn't be convinced that without corrections, LTCM would happen again—bigger and with more ramifications." Today, Bear, Lehman, and AIG have untold amounts of outlandish derivatives on their books. It could be years before anyone untangles what they're worth.
One other legacy of LTCM is "moral hazard": the prospect that other financial actors would take greater risks because at some level they'd assume that they, too, would be considered "too big to fail." Surely one can surmise that Fannie Mae and Freddie Mac overstepped in part because of an implied federal safety net that turned out to be a very real one.
Edward S. Lampert, the hedge fund tycoon who controls Sears Holdings (SHLD), worries about yet another twist. He says the current wave of federal intervention sends the opposite signal from what's intended: that officials are panicking because of broader instability. "As an investor, that was my immediate reaction" to the Fannie and Freddie moves, he says. "They completely destroyed confidence in any financial institution."
Lampert frets that with investment banks failing and merging, the resulting consolidation will concentrate risk and invite more rescues. "You are going to have Citi, JPMorgan, and Bank of America with $2 trillion-plus in assets each," he notes. "That's three times the size of Fannie and Freddie. Now if they end up with problems, what do you think is going to happen? They are too big to fail."
Pawel Dawidek first ported ZFS to FreeBSD from OpenSolaris in April of 2007. He continues to actively port new ZFS features from OpenSolaris, and focuses on improving overall ZFS stability. During the introduction to his talk at BSDCan, he explained that his goal was to offer an accessible view of ZFS internals. His discussion was broken into three sections, a review of the layers ZFS is built from and how they work together, a look at unique features found in ZFS and how they work internally, and a report on the current status of ZFS in FreeBSD.
The BSDCan website notes that Pawel is a FreeBSD committer, adding:
"In the FreeBSD project, he works mostly in the storage subsystems area (GEOM, file systems), security (disk encryption, opencrypto framework, IPsec, jails), but his code is also in many other parts of the system. Pawel currently lives in Warsaw, Poland, running his small company."
Derived from notes taken at a one-hour BSDCan talk by Pawel Dawidek, titled, A closer look at the ZFS file system. Simple administration, transactional semantics, end-to-end data integrity.
In a series of slides titled "ZFS, the internals", Pawel started with a diagram illustrating the many layers of ZFS, offering a quick overview of how it all fits together, and how it fits into FreeBSD. He then quickly moved from layer to layer.
zpool status -v which shows all errors as well and lists all files affected by these errors. An example use of this Pawel pointed out was that's it's easy to quickly determine exactly which files need to be restored from a backup. A feature found in the latest ZFS release, which Pawel is actively porting to FreeBSD, is the ability to use an entire device for caching, which he noted was similar to an L2 cache.
fsck. When data is modified, the change version is written to a new place on the disk rather than overwriting the old copy of the data. Once written, the pointers are updated to point to the new data. If there's a crash in the middle of an operation, the old pointers will still lead to the old data which will remain consistent and unmodified. Traversing the live filesystem is not easy when you have multiple datasets mounted, a feature provided by this layer. This allows you to synchronize mirrors, and is used when verifying all checksums in your pool.
/dev/zfs, the communication gate between userland tools such as zfs(8) and zpool(8) and ZFS, used to configure the kernel and modify ZFS pools. ZFS Features
Pawel described RAID-Z as "similar to RAID-5, and yet so much different". RAID-Z gains from the fact that ZFS uses copy on write, and never overwrites data, avoiding the above limitations with RAID-5.
RAID-Z is also self healing, because a checksum is written when data is written with RAID-Z, and then each time data is read the checksum is always validated. If the checksum doesn't validate, ZFS automatically attempts to reconstruct the data from the parity information, then validates this reconstructed data - if valid, it writes the corrected data back to the disk.
Another advantage to RAID-Z is that when a disk is replaced, it doesn't blindly copy the entire disk. Instead, it only copies actual data, so if a pool is almost empty synchronization can happen very quickly.
He then discussed hardware that does checksumming in the controller. For example, disks might be formatted with 520 byte sectors rather than 512 byte sectors, and the extra 8 bytes is then used to store checksum data. Pawel pointed out that this still does not provide end to end integrity, and can still be corrupted by a bad cable, in memory, or even by a buggy driver. Returning to the mail carrier analogy, he suggested they'd be saying something like: "We can only guarantee that when the package left our office, it was okay."
Other filesystems offer checksums providing block consistency verification, checking the block itself but not guaranteeing that the block is in the right place. Thus, a controller bug could mistakingly send writes to the wrong place, or phantom writes can happen when you think you wrote data but you didn't. Continuing the the mail carrier analogy, he offered: "Here is a package. It's not broken, but it may not be yours."
And then finally he looked at how every block is verified against an independent checksum in ZFS. Pointers are stored to the block in another block along with a checksum. When data is read, it can verify the data and that it really is the block being asked for. Stepping back, he noted that as data is stored in a tree, you have checksums going all the way up to the topmost block which offers a single checksum of all blocks in the filesystem. He described this global checksum as a cryptographically strong signature of the entire pool.
To maintain snapshots, ZFS tracks when a block was stored using a counter incremented each time an operation is written to disk, as well as a pointer to the block and a checksum. Every snapshot maintains its own dead block list, which is reviewed when a snapshot is destroyed, freeing blocks that meet the following conditions: they were born after the previous snapshot, born before the destroyed snapshot, they died after this snapshot was created, and they died before the next snapshot was created.
This synchronization happens from the top of the tree and works its way down, so if it is stopped mid-process by a crash, it is possible to pick up where it left off, or to obtain at least some of the data from the partially synchronized disk.
ZFS Status in FreeBSD
Pawel explained that he has already ported the most recent version of ZFS from OpenSolaris, and that it currently lives in his private Perforce source code repository. He noted that this port is completed code wise and everything works, but that he's working on writing regression tests. He's already written 2,000 tests, but these only cover half of ZFS functionality -- an illustration of just how many features ZFS has. The new code will not be comitted until he completes the writing of his regression tests, so he suggests "be patient".
Cool New Features in the Latest Port
When Will ZFS Be Production Ready?
Pawel notes that he's heard this question a lot. "The experimental status is very inconvenient," he commented to lots of laughter from the crowded room. He noted that he's currently the only maintainer, and suggested until someone comes along to co-maintain the code to help debug things when the filesystem gains more users he wouldn't be marking the code as production ready. He also commented that nobody has stepped up yet to co-maintain the code, so he expect is will be a while yet.
He went on to note that he's personally used ZFS on FreeBSD in production for 2 years, and on his laptop for more than 1 year, "it just works, and it doesn't lose data. It doesn't corrupt data, and you don't have to wait for fsck."
Questions and Answers
With this, Pawel opened the floor to questions.
Q:
A: Not yet. The regression tests are being written first, then the patch will be published, then it wil go into CVS.
Q: Will the new version of ZFS be able to talk to partitions created with the old version of ZFS?.
A: Yes, but you will need to use a command to update the volume if you want to access the new ZFS features.
Q: How does ZFS handle bad sectors on the disk
A: This can be handled by mirror disks or using RAID-Z. In addition, ZFS always replicates its metadata, and it's possible to configure it to also replicate data on a single disk.
Q: Does it support ACLs?
A: The new version does. In OpenSolaris they use filesystem attributres. In FreeBSD we use extended attributes. In the new version the two can be translated. It's also possible to implement POSIX ACLs, but this isn't likely to happen as it would make ZFS on FreeBSD incompatible with ZFS on OpenSolaris. There's also a Google Summer of Code project related to this.
Q: How does ZFS work with 64-bit architectures?
A: Another nice ZFS feature is that it has no endian dependencies. ZFS always writes in the architecture's endianness, and doesn't slow down writes by translating. When reading, it simply checks the order in which data was stored, then feeds bytes appropriately.
Q: Can you dynamically expand filesystems?
A: Yes.
Pawel then popped up a terminal and offered a live demonstration of how it works.
Q: How much space is allocated for snapshots?
A: No space is allocated for a snapshot until you start modifying it, then it allocates space as the filesystem changes.
Intrepid coder Bart Teeuwisse has written up an excellent technical account of creating "Tweet", a beautifully designed SearchMonkey app for Twitter. From a performance standpoint, writing a Twitter SearchMonkey app is particularly challenging, as Bart explains:
It turns out that execution speed of a SearchMonkey is key. To make the SearchMonkey Gallery a presentation monkey such as Tweet has to complete within a fraction of a second. Any call to fetch 3rd party takes too long to satisfy this requirement. Certainly calling Twitter's API whose fluctuating response times are all over the map.
Secondly, Twitter's profile API call takes a user ID, which first has to be extracted from Yahoo!'s indexed data. An additional data SearchMonkey can do that and whose output is the input to Tweet's profile feching data monkey. However, this chaining of data monkeys makes Tweet only slower.
Fortunately, Bart hit on a really clever solution: a mashup with Google App Engine, which acts as a simple proxy cache for Twitter data, which SearchMonkey can then consume. The result (after also adding Bart's own FriendNet infobar app):
Not only is the caching a nifty way to smooth out the API response times, but it also helps reduce the number of (rate-limited) API calls required. Read more about it at Bart's place.
Yahoo! 360?? - Dawn Patrol - Tweet, a Yahoo! SearchMonkey application to enhance Twitter user profilesTweet is a plugin for Yahoo! Search. Such plugins are called SearchMonkeys in honor of Greasemonkey for FireFox browser. Like Greasemonkey, SearchMonkey allows developers to enhance the experience, the search experience in this case. SearchMonkeys can enhance presentation with images and additional links or by combining Yahoo!'s Search index with other structured data.
Yahoo! Search users can add SearchMonkey applications to their profile on an opt-in basis. Add Tweet to yours if you like to get much improved search results for Twitter user profiles.
While Twitter user profiles are being indexed by all major Search engines, their summary is extremely poor. Google and Yahoo's results are nearly identical. Yahoo!'s summary of my profile (below) doesn't even include my full name (Bart Teeuwisse), which is on the page.
Luckily with SearchMonkey you can replace standard summaries with enhanced summaries. To improve Yahoo!'s Twitter user profile search results I wrote a SearchMonkey application called tweet that is triggered for all URLs matching *.twitter.com/*. Tweet calls Twitter's API to fetch user profile information not in the Yahoo! Search index. The result is a rich overview of a Twitter user, including last message (aka. tweet).
Sounds simple doesn't it? Contact Twitter's API, get profile, present profile. The SearchMonkey's architecture splits this into 2 monkeys:
Well, not quite.
It turns out that execution speed of a SearchMonkey is key. To make the SearchMonkey Gallery a presentation monkey such as Tweet has to complete within a fraction of a second. Any call to fetch 3rd party takes too long to satisfy this requirement. Certainly calling Twitter's API whose fluctuating response times are all over the map.
Secondly, Twitter's profile API call takes a user ID, which first has to be extracted from Yahoo!'s indexed data. An additional data SearchMonkey can do that and whose output is the input to Tweet's profile feching data monkey. However, this chaining of data monkeys makes Tweet only slower.
Thirdly -as I mentioned earlier- Twitter's API has wildly varying response times. And is by no means predictable enough to guarantee a prompt response. Furthermore Twitter is having scaling issues already. Adding a SearchMonkey that calls Twitter's API for up to 10 search results for each query could make things should Tweet gain many opt-in users.
Perhaps caching can help? The SearchMonkey platform does has some caching. Unfortunately SearchMonkey developers have no control over SearchMonkey's cache. Emperical data suggests that SearchMonkeys are cached for only a few minutes. Tweet could be cached much longer without sacrificing functionality.
To mitigate these challenges I decided to use a proxy of my own in between SearchMonkey & Twitter.
I first turned to Yahoo! Pipes, but Pipes doesn't give me caching control and the only XML output format is RSS not DataRSS. So I turned to Google's App Engine instead, which satisfies all my requirements. It offers Memcache caching, is build to scale, allows me to extract the Twitter user ID, make the Twitter API call and transform its response to DataRSS.
Even though this is my 1st Python application worth mentioning, I didn't have too much trouble writing it. App Engine's documentation combined with Python's tutorials were sufficient to answer my questions. The biggest obstacle I encountered is the lack of good XML/XLT libraries for Python. There isn't a clear winner to begin with and App Engine's restriction to pure Python libraries eliminates all candidates, as I learned the hard way.
I really like the Googel App Engine SDK. No hassle configuring a web server or data base. No need to be online even. I developped about half the proxy while vanpooling to and from work!
My proxy takes the URL of the search result as input from SearchMonkey. Given the trigger URL pattern these are all URLs to *.twitter.com. E.g. twitter.com, explore.twitter.com or m.twitter.com. The proxy first extracts the Twitter user ID, if any. In Twitter's URL schema, user IDs are the 1st part of the URLs path. E.g. bartt in twitter.com/bartt or twitter.com/bartt/friends
It then checks the Memcache for a profile for this ID. If it has one it composes the DataRSS response and exits. If it doesn't it calls Twitter's API. Succesfull API calls are parsed and stored in Memcache for -currently- 2 hours, before composing a DataRSS response. Failed calls return an empty DatRSS response.
My proxy speeds up cached profiles by a factor 3x to 10x. Most of the time, that is. Despite App Engine's claim to scale, it does have performance issues from time to time. App Engine had an outage for a day while I tested my proxy for example.
Odly enough, Twitter's API holds the record of the fastest response time, yet its average is many times App Engine's average response time (for cached profiles). App Engine's response time is very stable - about 200 milliseconds round trip from a west coast data center.
This doesn't make Tweet fast enough to be included into the SearchMonkey Gallery though. Not only is the proxy not fast enough, to that you'll have to add the XSLT process and 'render' times by SearchMonkey. Still, Tweet is now eminently more usable and shields Twitter from API overload.
Combine Tweet with FriendNet, one of my other monkeys for an even richer search result. FrienNet displays profiles and contacts embedded in the page. It combines hCard profiles with XFN links embedded on the page to present a social graph.
In collapsed mode -the default- FriendNet shows the number of profiles, cards and contacts found on the page by Yahoo! Search.
Expanded, FriendNet shows details of Twitter friends.
Got your own ideas for improving Yahoo! Search? Start monkeying around! You find everything you need at SearchMonkey on the Yahoo Developer Network.
Check out the SearchMonkey Gallery for more monkeys you can use. Or take my Better Amazon monkey for a spin.
Today, Yahoo! Search is taking another step in extending the Yahoo! Open Strategy with the launch of Yahoo! Search BOSS, a web services platform that allows developers and companies to create and launch web-scale search products by utilizing the same infrastructure and technology that powers Yahoo! Search.
Our goal with BOSS (Build your Own Search Service) is simple ? foster innovation in the search landscape. As anyone who follows the search industry knows, the barriers to successfully building a high quality, web-scale search engine are incredibly high. Doing so requires hundreds of millions of dollars of investment in engineering, sciences and core infrastructure ? from crawling and indexing technology to relevancy and machine learning algorithms, to stuff as mundane as data centers, servers and power. Because competing successfully in web search requires an investment of this scale, new players have effectively been prohibited from delivering credible alternatives to Yahoo! and Google. We believe the BOSS platform will begin to change that.
So what is BOSS?
BOSS is a new, open platform that offers programmatic access to the entire Yahoo! Search index via an API. BOSS allows developers to take advantage of Yahoo!��s production search infrastructure and technology, combine that with their own unique assets, and create their own search experiences. While search APIs have been available for some time, BOSS removes many of the usage restrictions that have prevented other companies from using them to build innovative new search engines.Here��s a quick summary of what��s available today:
? Ability to re-rank and blend results ? BOSS partners can re-rank search results as they see fit and blend Yahoo!��s results with proprietary and other web content in a single search experience
? Total flexibility on presentation ? Freedom to present search results using any user interface paradigm, without Yahoo! branding or attribution requirements
? BOSS Mashup Framework ? We��re releasing a Python library and UI templates that allow developers to easily mashup BOSS search results with other public data sources
? Web, news and image search ? At launch, developers will have access to web, news and image search and we��ll be adding more verticals soon
? Unlimited queries ? There are no rate limits on the number of queries per day
These capabilities are really just a first step ? we��re already working on expanding the API functionality and providing more access to Yahoo! Search Technology.
In addition to a self-serve API, we��re also partnering with a handful of Internet companies with large user bases or unique assets to collaboratively develop next gen search products using Yahoo!��s full suite of search technology. To learn more about BOSS Custom, click here.
What��s in it for Yahoo! and partners?
Why would Yahoo! open up its search infrastructure and technology to developers, entrepreneurs and companies who could use it to compete with us? It��s really quite simple. First, we believe that being open is core to Yahoo!��s future success ? opening our network, opening our own search experience via SearchMonkey, and now opening our search infrastructure via BOSS ? will lead to innovation both on Yahoo! and powered by Yahoo!. For BOSS, we see a virtuous circle in which partners deliver innovative search experiences, and as they grow their audiences and usage we have more data that can be used to improve our own Yahoo! Search experience and as a result, improve the quality of results our BOSS partners and their users get. Second, we do see new revenue streams from BOSS. In the coming months, we��ll be launching a monetization platform for BOSS that will enable Yahoo! to expand its ad network and enable BOSS partners to jointly participate in the compelling economics of search.What��s in it for users?
More choice. BOSS will enable a range of fundamentally different search experiences. These new search products will provide value to users along multiple dimensions, such as vertical specialization, new relevance indicators and ranking models, and innovative UI implementations. Our hope is that the resulting expansion in user choice will have the effect of fragmenting the increasingly consolidated search market in much the same way that cable TV dramatically increased programming choices for television viewers.Kick the tires and get started
Want to kick the tires on what BOSS-powered search could look like? As part of an alpha program, we��ve been working with a handful of start-ups and developers who have already begun using BOSS. Here are a few early examples of what��s possible with BOSS:? Me.dium, a start-up that��s built an innovative collaborative browsing product used BOSS to build a web-scale search engine that leverages its real-time surfing data. By combining the depth of the Yahoo! Search index with its insight into where users are browsing, Me.dium can provide its users with a unique buzz-based search experience.
? Hakia, a semantic search start-up, is using BOSS to access the Yahoo! Search index and dramatically increase the speed with which it can semantically analyze the web. With BOSS providing this important infrastructure, Hakia is able to deliver a language search experience that isn��t available from any of the ��big three�� search providers or other semantic search engines.
? Daylife To-Go is a new self-service, hosted publishing platform from Daylife. Anyone can use this platform to generate customizable pages and widgets. Daylife To-Go uses the BOSS API platform to power its web search module.
? Cluuz, a next-generation search engine prototype, generates easier-to-understand search results through semantic cluster graphs, image extraction and tag clouds. The Cluuz analysis is performed in real-time on results returned from the BOSS API.
To learn more about BOSS and get started using the API, visit the YDN. BOSS is open to all ? so check out the documentation, get a BOSS app ID, and start building the next generation of search.
Problem Statement
At 3rd&Urban, and in particular amp.fm (3rd&Urban is the parent company of amp.fm), our entire platform is built on top of Amazon Web Services products such as EC2, S3, and SimpleDB and driven by community-created content and interaction. Due to the nature of computer hardware -- especially those with moving parts -- while complete failure of an entire system is unlikely, failure of individual components within that system such as power sources and supplies, network cards, switches and routers, hard drives, processors, memory, and other components with an understood life expectancy is considered normal, if not rare, behavior. However, failure of any given component which results in outages which have crippling effects on the continued operations of the entire system are considered catastrophic. Designing and building fault-tolerance into any given system is critical to ensure that you always have back-up components in place to fall back on during an outage or failure of any given system component. Like any other data and community-centric company, we are committed to reducing the chance of a catastrophic system-wide failure to as close to zero as can be considered reasonable given understood component failure rates and unforeseen catastrophic events such as natural disasters.
While EC2 facilitates the ability to both add and replace instances on the fly, during the failure of an instance, at present time, any data on these instances that is not properly backed up will be lost. While backing up data to S3 is standard practice, data backups do not guarantee uninterrupted read/write access to that data, only the ability to recover from catastrophic failure, a process which, depending on the size of the data set, can take anywhere from a few minutes to a few hours to rebuild the effected data components. This time frame can potentially be even longer for data sets of considerable size and data structure complexity. As it relates to maintaining an always on, always accessible web business, we consider this a completely unacceptable scenario to potentially find ourselves faced with. As such, at the center of our system architecture resides a foundation of fault-tolerance techniques designed to ensure data persistence, redundancy, network accessibility, and automatic fail-over which, when combined together with off-the-shelf, open source software components, provides reasonable assurance of maintaining close-to-100% system up-time regardless of the failure of individual system components.Solution Summary
Amazon Web Services recently announced they are actively working on providing persistent storage as part of their EC2 offering, aiming to launch this service later this year. From the previously linked EC2 forum entry the Amazon EC2 team provides the reasoning behind this pre-beta release announcement,
"Many of you have been requesting that we let you know ahead of time about features that are currently under development so that you can better plan for how that functionality might integrate with your applications. To that end, we would like to share some details about a major upcoming feature that many of you have requested - persistent storage for EC2."
Speaking directly to,
"... so that you can better plan for how that functionality might integrate with your applications..."
... the primary focus of this paper is to present both a detailed overview as well as a working code base that will enable you to begin designing, building, testing, and deploying your EC2-based applications using a generalized persistent storage foundation, doing so today in both lieu of and in preparation for release of Amazon Web Services offering in this same space.
PLEASE NOTE: I have used generalized assumptions related to persistent storage solutions during the writing of this paper. Some of these assumptions extend from information that has been made public by AWS. I'll provide a summary of both the official announcement as well as Jeff Barr's (AWS Technical Evangelist) blog entry related to their persistent storage offering in the section that follows.
DISCLAIMER: There is no known direct or indirect connection between the material presented in this paper and the AWS persistent storage solution. While there is no reason to believe the same generalized ideas and technologies contained in this paper will be incompatible with Amazon's persistent storage offering when it becomes publicly available later this year, there is no guarantee this will be the case. While designing, building, testing, and deploying applications using the methodologies outlined in this paper, please do so with the understanding that you may have to re-design, re-build, re-test, and re-deploy certain aspects of (this|these) application(s) to take full advantage of the features and functionality provided by the public release of Amazon's persistent storage solution.
Please keep in mind, however, that regardless of any extended features and/or functionality introduced as part of the Amazon's public persistent storage release, the technologies and techniques describe in this paper will continue to work standalone, as-is.
The Solution
To ensure a proper understanding of what this solution provides and what it does not provide, the following two sections are a comparison of the publicly announced features of Amazon's persistent storage solution,
Features This Solution Provides
The following features are provided as part of this solution,
- Data redundancy via near-real-time synchronization of two block devices contained on two separate EC2 nodes using DRBD.
- Network mountable shares (NFS ) which provides the ability to mount these shares on more than one EC2 node at a time.
- Automatic fail-over between the primary and secondary DRBD nodes.
- Automatic and transparent remapping and remounting of an NFS share during the fail-over process.
- The ability to create snapshots of your volumes and back them up to Amazon S3.
- The ability to increase or reduce the size of any given volume that is part of the configuration, limited only by disk availability and capacity.
- At present time disk availability refers to the additional ephemeral block devices contained on m1.large and m1.x-large instance types.
- As already specified, while there are no guarantees this will work, in theory it will be possible to extend a logical volume with additional EC2 persistent storage block devices when this service becomes available.
Features This Solution Does NOT Provide
The following features are NOT provided as part of this solution,
- Highly durable persistent storage block devices that live independently of any given EC2 instance.
- The ability to create volumes ranging in size from 1 GB to 1 TB.
- Using LVM, it is possible to create logical volumes that range from 1k to the maximum capacity of your available ephemeral block devices.
- It's not possible, however, to extend things past the maximum size of the available ephemeral block devices.
The ability to attach and detach any given block device to and from any given EC2 instance.
- However, mounting block devices over NFS on multiple nodes does provide some of the benefits of this announced feature.