Wednesday, April 18, 2012

Introducing the Dental Microwear Image Library

Dental microwear, seen in the tiny pits and scratches on a tooth, provides lots of detailed data for inferring diet and chewing behavior in animals. Analyses are often conducted by digitizing highly magnified images of the tooth surface and counting up and classifying the various microscopic features. Animals with a certain percentage of pits and scratches may have browsing habits, whereas those with another profile may be grazers. By measuring extant animals with known diet, we can (hopefully) infer the diet of extinct animals.

Dental microwear; modified from figure 3 in Mihlbachler et al., 2012
In this age of increasingly open science, microwear studies can be problematic. A cornerstone of science is reproducibility - yet, inter-observer variation and error can greatly affect measured data. Furthermore, one study alone may generate dozens or hundreds of images. Even if you wanted to re-analyze teeth, it's pretty tough - how could you get access to the necessary images? Ideally, we want a world where anyone can access the raw image data, make their own observations, double-check published analyses, and add new data for comparison.

Thus, a new project - called the Dental Microwear Image Library, or DMIL - may change things. Assembled by Brian Lee Beatty and Matthew Mihlbachler, the website aims to become a clearinghouse for dental microwear images. This will allow greater standardization of analyses and hopefully better interpretations of paleoecology and diet for extinct organisms and modern organisms. The first data (from a recent paper in Paleontologia Electronica) are now posted, along with many other data sets.

Brian Lee Beatty (who blogs at The Aquatic Amniote and tweets as @Vanderhoofius) was kind enough to answer a few questions about the DMIL. Thanks, Brian!

Was there a particular moment or incident that inspired you to build the DMIL? If so, what was it?
As we set out to test and develop the method that Nikos Solounias and Gina Semprebon started, we found ourselves frustrated by not only the lack of information on methods that were given in most microwear papers, but also the inability for people to check their work. Interobserver error is a major cause of problems for microwear, and the only way for anyone to be aware of those differences is if they compare interpretations of microwear surfaces, not just their numbers on a spreadsheet. The DMIL was the only possible solution to the need to share such images.

How has community response been so far? Is there any particular type of skepticism that you're working to overcome?
The DMIL hasn't yet come up against skepticism, but our first paper on this method that uses it has.

What license, if any, are the data housed under? Or is it on a case-by-case basis?
There is no license for the data. We want it to be completely open-access and simply available.

How would you envision the DMIL 10 years from now? What goals might you have for the long-term?
We hope it will be a place that people can use to learn how to use the methods we are continuing to develop. I most sincerely hope that it will not only be home to our own data, but also be a place for others to deposit their data using similar methods so that more work of this sort is available in a similar, comparable format.
Authors of the recent paper in PE, along with a research assistant. Photos courtesy of Brian Lee Beatty.
For more information, check out the DMIL, or read the recent paper (open access) about the work.

Citation:
Mihlbachler, Matthew C., Beatty, Brian L., Caldera-Siu, Angela, Chan, Doris, and Lee, Richard, 2012. Error rates and observer bias in dental microwear analysis using light microscopy. Palaeontologia Electronica Vol. 15, Issue 1;12A,22p.

Thursday, April 5, 2012

Open Access in the UK - Comment Now!

The Research Councils UK (an umbrella organization overseeing much of the public scientific funding in that country, as well as funding for the arts and other worthy ventures) is soliciting comments on a new open access policy [PDF]. No matter what your opinion on open access, please comment. Mike Taylor, writing at SV-POW!, has further information and instructions.

Even if you don't live in the UK, it is worth letting the Research Councils know how you feel about the policy. Why? Because science (and scientific publishing) is inherently an international endeavo(u)r. I collaborate with colleagues in the UK all of the time, and many of the best papers I read these days have their origin across the pond. But, as with most scientific literature, access sometimes ain't easy. A more open scientific literature helps all of us, and each accessible paper raises the country's profile in the scientific community. Funding agencies always want more bang for their buck (or pound), and improving accessibility is one great way to do that.

So, drop a line to communications@rcuk.ac.uk by April 10 and use the subject "Open Access Feedback." Even a short sentence of support will do. Or three short sentences, as I did (basically using the argument in the paragraph above). Make your voice heard!

Wednesday, March 14, 2012

Curators: not just for museums anymore?

"The promise of the Internet-as-Alexandria is more than the rolling plenitude of information. It’s the ability of individuals to choreograph that information in idiosyncratic ways, the hope that individuals might feel invited by the gravitational pull of a broad and open commons to “rip, mix, and burn” — to curate." — Gideon Lewis-Kraus, 2007 [emphasis mine; paywall to original quote]

The internet is a very big place, and it can be exceptionally tough to keep track of everything that's of interest. A typical reader randomly browsing a topic ends up with a fair number of dead-end clicks—articles that just aren't that relevant or important. Fortunately, some individuals out there collate the best stuff, remix it, and push it to the outside world through blogs, news sites, Twitter, Facebook, Google+, and other venues. In a formal sense, these individuals are often termed "web curators" or "content curators".

In a nutshell, the web curator is not necessarily a content creator, but a content editor. I mean "editor" in the broad sense, of course—someone who selects interesting pieces and places them with other interesting pieces in new and meaningful ways. This is similar to what editors of magazines or anthologies do.

Museum Curators
Of course, the term "curator" was around well before the internet -- most notably in the museum profession (Wikipedia provides a pretty good summary, as does the US Bureau of Labor Statistics). In fact, my official job title is curator, so I can speak from some personal, professional experience. What exactly do curators do, then?
  • Direct the overall collection strategies for an institution. What to collect, what to deaccession, what to devote resources towards, etc.
  • Ensure the long-term survival of the collections.
  • Engage in original research (often using museum collections).
  • Present the collections to a broad audience, often through physical exhibits but also through various other media.
Depending on the type of institution and the field of study, a curator's duties may vary. For instance, a curator at an art museum may have slightly different tasks from that at a natural history museum, and collections managers may do some of the routine maintenance and preservation stuff at large museums. In any case, curation is a complex job.

Note that there is some overlap between the goals of a typical web curator and a museum curator. Both select and present collections of objects (fossils, or artwork, or blog posts) to a broad audience, but here the resemblance ends.

The British Museum - where curators reside. Image by awv, cc-by-2.0.

Why So Annoyed?
The contemporary usage of "web curator" is fundamentally misleading, at least judging from the above job description of a museum curator. Web curators collect, winnow, repackage, and disseminate information; they may have little role in content creation, and often have no concern for content longevity or archival. By the very nature of the internet, a web curator's work may be ephemeral (but not always). I would argue that this typical absence of the long view is a fundamental difference between most web curators and most museum curators.

Usage of the term "web curator" also muddies the waters around genuine digital curators. These are archivists, preservationists, and conservationists who work to ensure our digital heritage will be extant for the long haul. Digital curators in the pure sense aren't just repackaging links; these individuals ensure that the linked content will be around in 200 years. Calling yourself a curator doesn't mean you are one (similar to how not all museums are really museums, and loose applications of the term paleontologist, no matter how well-intentioned).

In part, I admit that some of my objections to the new usage of curator are a knee-jerk turf defense. I paid my dues, got my Ph.D., have an office in a museum basement. . .what have these internet upstarts done? I recognize this, and realize that such feelings are somewhat irrational. The English language changes constantly, and old words are often repurposed. After all, the web used to be just a product of a spider's backside. I just have to deal, right? On some level yes, but it still doesn't mean I have to like it! Nor does it mean I'm wrong.

The Most Important Objection
Admitting that definitions expand and contract, the fundamental issue here is that the phrase "web curator" is still basically meaningless. It serves to obfuscate, implying some kind of profundity where there may be none. In short, "web curator" is a buzzword.

A buzzword is corporate-speak that gussies up an otherwise mundane concept and makes it more intriguing (and profitable). Consider some examples. Value-added. Holistic. Accountability. All perfectly nice terms whose vague usage gives me a headache.

The big problem here—and a key quality that makes "curator" such a great buzzword—is that most people have no clue what a curator does. At best, folks have some vague notion of a curator as a person with a fancy degree and hipster glasses who hangs paintings on a wall and maybe writes some label copy. "Web curation" fits this stereotype and thus is a masterfully empty use of an important-sounding term (see these links for some choice, typical usages).

A solution?
I'm a big fan of calling a spade a spade. "Web curators" provide a valuable service, but the title unfortunately misleads. Just read the phrase "Real-time curators need to add participation widgets", and tell me it's not slightly silly! Were this statement not from a rather well-known blogger, the phrase could just as easily have originated in the Web Economy Bullsh*t Generator. In fact, the cited example is the perfect storm* of all that is wrong with buzzword-led thinking. 

Mike Taylor has pressed me on my objections to the term "web curator," asking for an alternative. I see nothing wrong with "editor". As outlined above, it's a much more accurate description of what "web curators" do. An editor is a skilled person who practices the art of identifying relevance and distributing the results. At its core, curated content on the web is part of a web anthology, just like an anthology of prose or poetry. The only difference is the digital format. Could we ever find a better, more descriptive term than "editor" for this role? In the digital realm, "curator" should be reserved for those who go beyond a primarily editorial role, to preservation, archival, and conservation.
 

So, ditch the web curator. Web editor, please.
-------------
*perfect storm = buzzword. Yes, I was being ironic by using it. Very meta, huh??
Thanks to Bora Zivkovic, Mike Taylor, Mike Keesey, Tori Herridge, and others for stimulating discussion and feedback that led to this post.

Tuesday, March 6, 2012

Self-archival: a good start, but not the full solution

We all want our work to be discovered, read and cited. There is little doubt that closed access systems hamper this - a paywall to an article is a hefty obstacle, and we all encounter them at least occasionally no matter how extensive our library access is. From an author's perspective, freely-available PDFs of their work are a major boost.

In recent discussions on Twitter and in the blogosphere, I've chatted with Mike Taylor, Ross Mounce, and others about self-archival as one of many mechanisms to bring about open access. Mike's recent blog post at SV-POW! summarizes much of the discussion to date, and I thank him for helping me to crystalize my thoughts on the topic.

For those who are not familiar with the term, self-archival refers to placing a freely-downloadable copy of a publication (or other work) on one's personal (or departmental, or whatever) web page. In this post, I want to discuss the pros and cons of such an approach.

Pros
  • The PDF is freely available to anyone who wants to see it. No paywalls. No hassle.
  • Once picked up by search engines, your posting may be the first one web users find - even above the "official" journal page!
  • If users browse your website with the PDF, it means that they might discover closely-related work. This can be a big plus for getting the word out about your research program. 
Cautions
  • A personal archive is probably not a permanent archive. Barring special arrangements, your personal or institutional web page is not likely to last substantively beyond your lifetime. Free hosting services such as WordPress may not be around in 20 years (remember Geocities?), so it may be worthwhile to pay for hosting. And make sure your descendents pay for hosting, or that your departmental web administrator doesn't delete your page 15 years after you retire. I have little faith that the PDFs I post on my own web page will be around 200 years from now, at least at that website. That sure would stink for that researcher in 2212, who wants to read all about ceratopsian sinuses.
  • Author-hosted archives are not independent. There is nothing to prevent someone from removing embarrassing details or adding fraudulent information to their publications, and little that a casual reader can do to detect such fraud. The great majority of academic authors are honest - it's that tiny minority we have to watch out for. An independent archive, hosted by an institution, library, or publisher, provides a firewall protecting the literature from the authors.
  • As article-level metrics gain prominence, author-hosted PDFs may skew some statistics. For instance, let's say I publish a paper in PLoS ONE, and also post a copy of the PDF to my site. Because PLoS ONE records and posts view and download statistics for its own site, any downloads or views from my site are not recorded there. Thus, the statistics are spread across several venues. This is not a major issue in my opinion, but some people may care.
  • Under the terms of publication, a publisher may not allow you to post a PDF of your paper. Or, they may only allow you to post a pre-review copy. Or a post-review, unformatted copy. Things get complicated quickly, especially for those concerned about following the letter of the law.
The Up-Shot
If you are active researcher, you should be posting whatever PDFs of your own work that you (legally) can.  If you don't, you're missing out on innumerable opportunities to publicize your work and interact with colleagues. However, personal archiving is not enough to ensure permanence. For the long-term, a bigger solution is needed. Institutional archives, journal archives, society archives, whatever. The ultimate answer may take some time to sort itself out.

    Friday, March 2, 2012

    The Open Museum Notebook - Torosaurus Style

    A new paper on the Torosaurus / Triceratops issue was just published in PLoS ONE, bringing some additional analysis to the table. I won't comment on it any more here (I'm saving my thoughts for a formal reply on the PLoS ONE website itself), other than to refer you to my own paper and the Scannella & Horner response.

    In any case, I have a pile of notes from my own work on Torosaurus (or whatever we should call it), and figured it was time I distribute them a little more widely. So, I just uploaded my notes on the Yale Torosaurus specimens to figshare.com. There isn't really anything earthshaking in there (most of the meat of it has been previously published), but in any case now other folks can use them. The sketches of real bone vs. reconstruction should be particularly useful.

    My sincere hope is that at least a few other paleontologists will follow suit with their own notebooks - there are a lot of unused data that will never see the light of day otherwise. I also have a goal of gradually digitizing and posting my other museum notebooks, but that will probably take some time!

    Citation and Link
    Notes and Observations on Specimens of Torosaurus at the Yale Peabody Museum of Natural History. Andrew Farke. Figshare. Retrieved 15:40, March 02, 2012 hdl.handle.net/10779/664bf2cb5ac486da32c7fb7261e595cd

    Update: Since this posted, I have uploaded a number of other notebooks. Find them on my figshare author page.

    Wednesday, February 8, 2012

    Restoring that sense of wonder

    These can be depressing times for a paleontologist - funding is poor for most, the job market is dim for many talented friends and colleagues, and rhetoric-ridden battles for scholarly publishing rage. That's enough to suck the joy right out of the field. In instances like this, it's nice to step back for a second and think about the really cool stuff going on.

    So, I've put together a list of wondrous things that have happened in paleontology over the past several years. Why are they cool to me? Mostly because they challenge ideas that I acquired while a little, dinosaur-obsessed kid. And they also challenge ideas I've acquired as an "educated" professional. Sometimes it's nice to have our comfort zone stretched.

    Symbols of the new paleontological revolution: an eye-catching Sinosauropteryx crouches on top of mammoth DNA, overlain on a thin-section of dinosaur bone (sources at end)
    In no particular order:
    • We know what colors were on parts of the body of some dinosaurs. Really. How cool is that? Sure, it's not perfect, and there is lots we'll never know, but the mere fact that you can plausibly reconstruct parts of the pelage of a feathered dinosaur is amazing. Especially because I had always believed the truism that we'd know the texture of dinosaur skins, but never the color.
    • I can download a genetic sequence from a woolly mammoth. Or a Neanderthal. Or any number of extinct organisms. I had always known that Jurassic Park would never be a reality. It probably never will be (at least for non-avian dinosaurs). But to stare at the A's, T's, G's, and C's of an extinct organism still gives me some goosebumps.
    • I can listen to a Jurassic katydid. Yes, yes, there are some assumptions in the reconstruction. But let's suspend criticism for a moment, and accept that it's probably at least a decent approximation. These are noises that haven't been heard in 165 million years.
    • We know the sex of some individual dinosaur specimens. Thanks to studies of medullary bone and comparative anatomy, the seemingly impossible is made real. Wow!
    • Similarly, we know the age of some dinosaur individuals at death (give or take a few years). The notion that sauropods only got big because they grew for a century can't be supported anymore. Once again - wow!

    This is just my personal list - what's on yours?

    Sources for image: Mammoth DNA sequence in background from GenBank Accession FJ655900 (published by Enk et al., 2009); dinosaur bone histological section modified from Woodward et al. 2011 Figure 1C (colors inverted and adjusted); Sinosauropteryx modified from original by Marty Martunuik. Image released under Creative Commons Attribution-Share Alike 3.0 Unported license.

    Tuesday, February 7, 2012

    How Big Commercial Publishers Can Help Themselves

    Big commercial publishers - especially Elsevier - have been getting a lot of flack lately. There's the usual background noise about high costs of institutional subscriptions and individual PDFs for non-subscribers, and now we have concerns over SOPA, PIPA, RWA and the burgeoning Elsevier boycott. I think it's fair to say that the argument has been dominated most strongly by the publishers' critics. Nonetheless, there is invariably someone who pipes up in comment threads (or in posts at sites like The Scholarly Kitchen) in defense of the publishers.

    Pro-commercial publisher arguments almost always include the term "added value" or something similar. In other words, the big publishers add something beyond the raw manuscript and figures that are provided by the authors. I think very few people will dispute this claim, at least at its face*. The publishers:
    • facilitate peer review by paying for a manuscript handling system (either licensing a commercial product or installing an open source product on servers they pay for) [note that this is not the same as doing the peer review, which is done by volunteer referees and unpaid or minimally-paid editors]
    • do some copy editing
    • format the manuscripts into a pretty PDF and web page
    • provide a veneer of respectability with well-known journal "brands"
    • distribute the journals to libraries and interested readers, via subscriptions, web hosting, and proprietary search engines
    • and other miscellaneous things
    [*To forestall the inevitable comments, yes, some of these "services" are of dubious value to many users]

    Look, I appreciate the fact that all of this costs money. Somebody needs to be paid to do the formatting into the appropriate medium (whether web page or PDF), technical staff need to make sure the authors submit the files in the right format, it costs money to run a server, programmers don't come cheap, and all of the various functions of a business/journal aren't free (office space, salaries for necessary employees, etc.).

    But does it really cost so much that publishers have to charge $37.95 for a single PDF file, or $392 for a personal subscription to a journal?

    Maybe the answer is yes (forgetting the 30%+ profits for many major publishers). Maybe it does cost a lot of money to produce an article. Fine. Just do a better job of convincing me that it's worth it. Particularly when some of the most labor-intensive tasks (typesetting and peer review) are provided for free by the authors and their colleagues.

    Many large publishers have an established list of things they do that cost money. They've done a decent job of publicizing these talking points, judging by the facts that they show up so often in comment feeds and that I was able to assemble the bullet points above virtually from memory.

    However, publishers have performed miserably at convincing us that $37.95 is a reasonable price for a PDF download. Elsevier and company could deflect much criticism if they were to be more honest and transparent about the costs behind a journal article. How much time/money actually goes into formatting? How much does it really cost to serve a file to the internet, over multiple years? What is the honest per-article cost for the manuscript submission system? How many people actually buy articles? Instead we're stuck with the broken record of "oh, this stuff costs money, OA advocates just think it all happens for free. . ."

    Finally, here's my most pressing question: If economies of scale apply to publishing, why are the largest publishers providing some of the most expensive services? (in terms of solo journal subscription rates, individual PDF downloads, and open access fees) Wow, would I love the answer to that one!

    Post script: It seems that many folks are having similar thoughts. Check out Björn Brembs' round-up here.