Showing posts with label controlled vocabulary. Show all posts
Showing posts with label controlled vocabulary. Show all posts

Wednesday, April 10, 2013

Sorting the Laundry and “CVS”




“CVS” is an abbreviation we have used previously for that cross-referencing ability provided in Interclipper and other database browse/search environments. CVS stands for “chinese, vegetable, spicy” and we use this as an example of how an index (or indexes), in this case a book of recipes, needs to be flexible and mulch-dimensional to optimize users browsing power.  In a recipe book, you would never ask "is the recipe Chinese OR vegetable OR spicy?" These represent different lenses through which you can organize many recipes, i.e., ethnicity, ingredient type, novel qualities of the food. But they are not mutually exclusive.

We often use CVS as shorthand for “multi-dimensional indexing”. It simply means that terms and groups of terms are organized in a coding framework/map that allows for different facets to be visuallized--not washed out by a single hierarchy or the alphabet.

Wednesday, February 13, 2013

Meaning Mapping for Ballparks, not Bulls eyes


In a number of our projects, we have had difficulty training our clients and partners NOT to get too specific by trying to imagine every possible future user while creating the controlled vocabularies and multi-dimensional indexes. (When we did this on an early project we ended up with an “out-of-control” controlled vocabulary.) The but the indexing process is not about naming things, rather sorting them into a collection of baskets of meaning we create. The structure of baskets can grow, change, expand or contract over time (an "iterative" process) and provide a lot of retrieval power and more than ample browsing power.

A misleading concept that Google reinforces in the digital age is the idea that accessing multi-media is about hitting bulls eyes or getting home runs. But one of the most underrated features of Google is not its powerful secret engine for retrieval that we will never understand, but the way it now leverages 10+ years of our “near misses”.  Google’s correlative database gives us a quick list of potential things we meant, but, like many others users, have misspelled or mis-Googled and eventually found. Even Google knows that searches are actually less about searching, and more about matching meaning to users’ desires. "Browsing" is still a mode of operating on the web and "Googling" is something different--more specific. We remember and embrace the promise of the web as a place to explore and browse...

Tuesday, December 4, 2012

Multi-Dimensional Indexing: A Dynamic Process


Traditional cataloging and indexing might typically be bounded in scope--like in the case of a collection catalogue or an index is developed for a particular book.  But oral history collections and other digital collections are often associated with active projects that are growing over time. If the indexing process is strongly content-informed, and the content is changing dynamically over time, then the indexing process must not only be an iterative one, but a dynamic one as well.

What does this mean for our controlled vocabulary development process? In order to begin to capture the breadth and diversity of a collection via an index, the more content used to contribute to its development, the better. But how much of the content should inform the index, and when? Is it okay not to reevaluate the earliest indexed material and code it up with the evolving framework?




Wednesday, November 28, 2012

Multi-Dimensional Indexing: An Iterative Process


Developing and refining a catalogue or index is an iterative process that includes cycles of brainstorming, organizing, arranging, testing, reorganizing, editing, and publication. Iterative means we expect to go back and adjust work done earlier in the process, informed by things we learn later in the process. This may seem inefficient if we are comparing to other types of work. But indexing is more like writing or music composition, where the goal is quality content and the product comes as a result of a number of drafts. Although some brilliant artists create and compose well spontaneously, others require multiple revisions of initial drafts. In indexing, we behave more like the latter artists, revising significantly our first attempts based on an evolving conception of what's important or feedback from an audience.

Thursday, August 30, 2012

Editor


As with any published document, an editor is needed to assure consistency and quality. An editorial role is also important to unify inconsistencies that may emerge between different annotators of the audio and video by dictating a style and giving feedback until the desired “voice” is achieved. The work may be as much managerial in nature as it is an editing role.  In general, the editor should be the master of all the text created in the annotation/indexingprocess. The editor would not be expected to have heard every minute of the original audio or video, but they would be expected to have read most or every word annotated within the collection.  In a larger project, the editor might delegate editorial tasks to trusted personnel, but ultimately there should be one person at the top of a hierarchy who is ultimately responsible for all published content.

For newly developed controlled vocabulary, there is also an editorial role associated with approving new terms. In the context of a custom/local controlled vocabulary being developed from scratch, the editorial process occurs as proposed terms are agreed upon and finalized.  In the case of updating, amending or expanding on an existing standard or local controlled vocabulary, content management systems like CONTENTdm have features that cue and allow a librarian to approve specific terms that have been added. In that case, the librarian who has the authority to approve or disapprove of a term being added is acting as an editor as well.

Thursday, August 23, 2012

Indexer/Coder


The role of indexer entails two parts of the indexing work. There is developing the index—i.e., brainstorming and organizing the customized controlled vocabulary to be used, and then actually applying the index to the content, which could be referred to as coding. As with all the other roles discussed here, this work may all be done by the same person. 

Developing an index, or more specifically the controlled vocabulary used as the index, may include contributions from anyone familiar with the content of the oral history interviews and the subject area in general. The annotator is typically best equipped to provide the most specific and topical input relative to the material that has been annotated. However, the controlled vocabulary is based not only on the very specific content of the collection but outside factors as well. Existing thesauri such as TGM and LCSH can be drawn upon to develop the controlled vocabulary. Also, the users--the anticipated audience of the collection—should be explicitly defined and considered when choosing terms (e.g., a local term for an object might be more appropriate than what LOC calls it.)  Developing a controlled vocabulary is particularly challenging in an on-going project, as the terms chosen and the architectural structure of terms (e.g., hierarchy) will necessarily change as more material is added. Ideally, the controlled vocabulary is developed after all interviewing and annotation is complete, though sometimes this is not possible. Developing the controlled vocabulary works best as a collaborative, iterative process (including drafting, debating, and test application) aimed at a comprehensive taxonomy that represents the whole collection.

Applying the controlled vocabulary is another task of the indexer, or more specifically the coder.  Coding is sometimes done by more than one person, and often done by someone different than the person who composed the annotation originally. Frequently, the coder applies the controlled vocabulary based on the annotation summary only, not by actually listening to the original recorded passage, which has certain advantages and disadvantages. The key challenge with this job is maintaining a reasonable amount of “intercoder reliability”—i.e., that two indexers/coders independently will assign the same vocabulary term to the same passage. Some inconsistency and subjectivity is expected in this process—just as two people would not index a book in exactly the same way. Ideally, one person, such as a lead indexer or the editor, should oversee the indexing/coding and develop some means of quality control of to assure a reasonable amount of consistency.

Return to Oral History Digital Indexing Roles.

Wednesday, February 15, 2012

Locative Metadata

When we index at Randforce, we develop controlled vocabularies (also known as thesauri) for annotated passages of audio and video. Technically, the objects we are indexing are metadata themselves. (Thus we create meta-metadata!) For discussion purposes, the objects we are indexing are a/v clips.

We often conceptualize the process of indexing to be less like "labeling" something (i.e., what is it called?) and more like "putting it somewhere" (i.e., where does it belong?) and we sometimes call this "sorting the laundry." The proverbial laundry baskets are created by us indexers and the objects influence its creation in a meaningful way. It seems to me, that when these terms are fed back to the object, they are a unique type of metadata.

I'd like to propose that this type of metadata being created might be called "locative metadata". Locative metadata, conceptually, is more than an attribute of the object. Locative metadata implies "where it is" relative to other objects in the collection, not just what it is about. Locative metadata might also be a purely digital concept exactly because an object can reside in more than one location at a time (without needing to take up additional space). In this sense, library subject headings--from the book's perspective--is locative metadata, as are hyperlinks to an object from the objects' perspective.


Wednesday, November 23, 2011

What is digital indexing?

“Digital Indexing” is shorthand we often use to describe our work in audio/video content management for oral history. Digital Indexing encompasses the processes and various software tools we employ to annotate and index a/v recordings electronically. Although we have no formal definition for digital indexing, the Illinois State Museum described it for their Audio-Video Barn site during our partnership under an IMLS leadership grant in 2009-2010:

“Digital Indexing is a method of defining starting and ending points to an audio or video and then describing that portion of media in a set of notes that can be searched through keywords and control words. The results of this method are a set of searchable notes and the ability to instantly watch the exact corresponding portion of audio and video.”

http://avbarn.museum.state.il.us/education/oralhistoryhowto/processing/indexing

Our work with “keywords and control words” will be the subject of a later post about “Multi-Dimensional Indexing”, which is our unique approach to controlled vocabulary in electronic environments. Future posts here will share some methods for summary-based annotation (typically applied before or in lieu of transcription) and an overview of The Interclipper™ software, a database system we find ideal for the challenge of recorded audio-video content management.

Phone: 800-554-1047 - E-mail: info@randforce.com
Web Site Copyright © 2011 The Randforce Associates, LLC