ESC
其他 10 分钟阅读

The Slow Formation of Durable Software

The Slow Formation of Durable Software

来源:Hacker News

Alongside another new Ph.D. who had studied the history of technology, Jim Sparrow (now a history professor at the University of Chicago), we built a website called ECHO: Exploring and Collecting History Online, funded by the Alfred P. Sloan Foundation. “Collecting” was a new gerund for the center, and hinted at an expansion of our methods. Since science was growing exponentially but the number of historians of science was not growing at all, we thought we could use the web to help scientists self-document their work. For this, our site needed to be interactive, so we started to tinker with “tools” — small web apps — that could not only display history online, but also allow it to be uploaded, sorted, and archived.

Toward this end, I wrote several applications in PHP, then swiftly becoming the standard programming language for web applications, because it commingled well with HTML, the web’s lingua franca. One of these apps, Web Scrapbook, got a bit of traction, especially in classes, since it allowed students to capture images, links, and other resources from the web browser into a collection that could be shared with their classmates and instructor. Users clicked on a browser bookmark, which had a small bit of JavaScript embedded in it, to select the desired items and pass the data to the application. Web Scrapbook was not, shall we say, a rigorous piece of software — my initial version sent passwords, unencrypted, across the internet — and it paid only basic attention to metadata and other scholarly information that would be needed to assist in writing an article or book, and to form a proper footnote.

Fortunately, at the same time, another colleague at the center, Elena Razlogova, who was the webmaster at RRCHNM while pursuing her Ph.D. (and is now a history professor at Concordia University), was working on a more scholarly application called Scribe, for taking notes and storing citations. Elena used a database application called FileMaker to build Scribe, which users downloaded and ran on their personal computers.

Scribe became popular among fellow historians as a good, free replacement for commercial software like EndNote. It ran locally on a Mac or PC, and had the organizational, search, metadata, and annotation capabilities that Zotero would later expand upon. Elena made several improvements to Scribe in the first few years of the century, as I worked in parallel on Web Scrapbook.

So by 2003, we knew how to write web apps and standalone apps, both helpful but both also naggingly insufficient. We were now starting most of our research on the web, and we thought that any robust research application of the future had to be connected to this environment where primary and secondary sources increasingly resided, either as metadata in library catalogs, or as full objects if they had been digitized. What we really needed was an application that had some aspects of Scribe and some aspects of Web Scrapbook: a full-fledged research tool that could exist independently of the web, but be fully cognizant of what was going on in the web browser, aware as the researcher surfed across digital collections and library catalogs.

On November 29, 2003, Roy asked Elena via email what she would like to do for a next version of Scribe. Roy, Tom Scheinfeldt, another recent history Ph.D. who had joined us to work on the September 11 Digital Archive (now a professor of digital humanities at the University of Connecticut), and I were working on a grant proposal for a new stage of ECHO, and we thought we would include some improved software tools as part of that. Elena wrote back on December 1, 2003:

Roy– Here’s what I’d like to do:

setup to connect to online bibliographic databases more citation styles (at least APA and legal) setup to add custom citation styles setup to add more fields and types of references user/password setup to share and co-edit online bibliographies/notes setup to work locally then publish all or parts of bibliographies/notes online (like ical) foreign language support better manual online discussion list (using the software marty had installed)

best, elena

This was a good list. Tom wrote back that “it would be nice to update and ‘webify’ Scribe.” A good word: webify. Elena, Tom, Roy, and I agreed on this essential next step for Scribe: we needed to find a way to put it into the web browser, like Web Scrapbook, while retaining Scribe’s provenance as a detail-oriented research tool. As I later summarized our holy grail in a journal article:

We wanted the best of both worlds: the best parts of standalone applications and Web applications. We envisioned a tool that lived in the browser and that was very smart about what was going on in the browser, including the recognition of scholarly metadata and objects, and that could interact with elements both on the desktop, such as word processors or other programs, and via standards, services, and communication protocols to other tools and resources across the Web.

But how to merge these disparate worlds? That was not at all clear in 2003.

At RRCHNM, we had a large whiteboard with vague ideas, some of them serious and some of them silly, visible to everyone, just in case it gave someone a spark. Over group lunches, the faculty and staff at the center would often talk about new tech we had heard about, and muse about whether we could apply that tech to historical research. Our collective was growing quickly. In the summer of 2004, we were joined by, among others, Josh Greenberg, who had a Ph.D. in science and technology studies (and is now a program director at the Alfred P. Sloan Foundation); Sharon Leon, who had a Ph.D. in American history (and is now the Co-CEO at Digital Scholar); and Simon Kornblith, a young teenager (and son of a historian and friend of Roy’s) who was truly gifted at software development and many other things. (Simon would go on to get a Ph.D. in Brain and Cognitive Sciences at MIT, write important AI papers with future Nobelist Geoffrey Hinton, and work at Anthropic.)

One project that caught our eye was a spinoff of the open-source Mozilla web browser that came to be known as Firefox. That summer Firefox was in beta, and it would be officially released later that year. The new browser had much to recommend it. The original Mozilla browser, descended from Netscape and the early days of the web, had, over time, become slow and bloated. Firefox, on the other hand, was fast and lightweight. But much more exciting was how Firefox embraced and foregrounded XUL, which sounds like an alien god from a 1950s sci-fi movie, but stands for XML User Interface Language. It dawned on us that with XUL, you could — sweet mercy — customize and extend the web browser in any way you wanted.

The day Firefox was released to the world on November 9, 2004, Josh and I — always the earliest of adopters — downloaded it and started to tinker with its possibilities. That night I wrote an email to Josh, Elena, Tom, and Roy with the subject line “Firefox + XUL + AWS = OpenScribe/Online Scribe?” The AWS I was talking about was not the future cloud computing service from Amazon, but an API Amazon provided to retrieve information about books. I thought we could use a combination of a Firefox extension written in XUL and services like Amazon’s to “automatically load citation info into the fields.” Josh wrote back just before midnight:

Please pardon my language when I say “Holy crap, that’s cool!”

The possibilities are astonishing - it looks like there are XPCOM wrappers for mySQL that would allow us to access a database (either remote or local), entirely replicating Scribe using a relational database (or, offering a custom “Scrapbook Browser.”) Alternatively, we could chuck the database and use a Firefox interface to navigate Scribe XML data directly. Regardless, immediate cross-platform compatibility.

Either way, we could use a combination of Amazon and Proquest/ISI Citation Index to make it much easier to create bibliographic objects, and at the least create a Mac version that would use either Cocoa or more basic scripting to allow citation from within a word processor…

Neat-o!

With an open-source database instead of FileMaker to store bibliographic data and notes, and with a Firefox extension of the proper composition, all of the ingredients were within reach to create the research tool of our dreams.

That winter, with Roy as the principal investigator and Josh and me as co-directors, we applied for a grant from the Institute of Museum and Library Services. The abstract from that grant, which we titled “SmartFox: the Scholar’s Browser for Digital Collections,” displayed our new level of clarity, although we still envisioned what we wanted to make as a “set of tools” — such as the item capture tool from Web Scrapbook and Scribe’s note-taking interface — rather than one app:

The web browser has become the primary means for accessing information, documents, and artifacts from libraries and museums around the country and the world, thanks in large part to the tremendous commitment these institutions have made to bringing their collections online (as either simple citations or complete text and images). Unfortunately for scholars, while tens of millions of dollars have been spent to create digital resources, far less funding and effort has been allocated for the development of tools to facilitate the use of these resources. The browser remains merely a passive window allowing one to view, but not easily collect, annotate, or manipulate these objects. Moreover, from the user’s perspective individual library and museum collections remain just that—separate websites with distinct designs and different ways of displaying their information, making traditional scholarly practices of bringing together and studying objects of interest from across these collections unnecessarily difficult.

SmartFox, a set of tools incorporated into popular, open, and free web software, will address these major problems by creating a web browser that is “smarter” in two key ways. First, one tool will enable the browser to intelligently sense when its user is viewing a digital library or museum object; this will allow the browser to capture information from the page automatically, such as the creator, title, date of creation, and copyright information. Second, another tool will store and organize this information, as well as full copies of items and web pages (not just their citation information) if so desired by the user and permitted by the institution’s site, allowing the user to sort, annotate, search, and manipulate these individualized collections created for scholarly purposes. Critically, all of this will occur within the web browser itself, not in a separate, standalone application; the web browser will be used not just to discover information, but also to collect, organize, and analyze scholarly materials.

(Zotero did indeed begin its life with a name, SmartFox, that was a riff on Firefox, and our potential trademark violation somehow got worse when we rebranded the app as Firefox Scholar in September 2005. More on that later.)

Even before IMLS funding had come through, Simon had started working with David Norton, another software developer, on a prototype of SmartFox. By July 2005, they had produced a very early “0.0.1” version, with the metadata panel on top, the notes field below, and the folders on the left; all of these would later be combined in a much better unified drawer.

Simon, who was paying attention not only to public releases of Firefox but also to the core code development itself, pointed out that summer that Mozilla planned to include “mozStorage” in Firefox, which could act as an interface to a database. (Firefox eventually included SQLite as its native, open-source database.)

Also in the productive summer of 2005, we hired Dan Stillman, a friend of an RRCHNM staff member, to help us with the center’s burgeoning set of servers and increasingly complex digital platforms. Dan fixed so many accumulated issues so well that by January 2006, Josh and I had asked him to join Simon and David on the dev team. (To this day, Dan is the lead developer of Zotero; what a run.) We also added Sean Takats, who had both a doctorate in French history and extensive technical experience, and became Zotero’s longstanding director in addition to his academic career in history and digital humanities; Kari Kraus, our technology evangelist, who became a professor in the College of Information Studies and the Department of English at the University of Maryland; and, later that year, Trevor Owens as our outreach coordinator. (Trevor is now the Chief Research Officer of the American Institute of Physics.)