mirrors/zotero - Ayakael: My personal forge

mirrors/zotero

Author	SHA1	Message	Date
Simon Kornblith	6305e4cada	closes #55 , export bibliography to printable version closes #4, Make printable version - moves functions for creating and deleting hidden browser objects to scholar.js (from ingester.js), since these are necessary for printing as well - allows saving bibliography in HTML or printing bibliography. style support is not yet complete (pending finalization of 0.9 version of CSL specification).	2006-07-27 23:01:55 +00:00
Simon Kornblith	c64e5c841f	closes #78 , figure out import/export architecture closes #100, migrate ingester to Scholar.Translate closes #88, migrate scrapers away from RDF closes #9, pull out LC subject heading tags references #87, add fromArray() and toArray() methods to item objects API changes: all translation (import/export/web) now goes through Scholar.Translate all Scholar-specific functions in scrapers start with "Scholar." rather than the jumbled up piggy bank un-namespaced confusion scrapers now longer specify items through RDF (the beginning of an item.fromArray()-like function exists in Scholar.Translate.prototype._itemDone()) scrapers can be any combination of import, export, and web (type is the sum of 1/2/4 respectively) scrapers now contain functions (doImport, doExport, doWeb) rather than loose code scrapers can call functions in other scrapers or just call the function to translate itself export accesses items item-by-item, rather than accepting a huge array of items MARC functions are now in the MARC import translator, and accessed by the web translators new features: import now works rudimentary RDF (unqualified dublin core only), RIS, and MARC import translators are implemented (although they are a little picky with respect to file extensions at the moment) items appear as they are scraped MARC import translator pulls out tags, although this seems to slow things down no icon appears next to a the URL when Scholar hasn't detected metadata, since this seemed somewhat confusing apologizes for the size of this diff. i figured if i was going to re-write the API, i might as well do it all at once and get everything working right.	2006-07-17 04:06:58 +00:00
Simon Kornblith	8b4a44be0f	fixes a bug that made the Google Books translator not appear adjusts the Google Books translator to work with the latest revision of the site renames the MODS translator to just MODS, because "Metadata Object Description Schema (MODS)" was too long for the export dialog	2006-06-30 19:21:36 +00:00
Simon Kornblith	77282c3edc	- fixes a bug that could result in scrapers using utilities.processDocuments malfunctioning - fixes a bug that could result in the Scrape Progress chrome thingy sticking around forever - makes chrome thingy disappear when URL changes or when tabs are switched	2006-06-29 03:22:10 +00:00
Simon Kornblith	45b9234996	addresses #78 , figure out import/export architecture - changes scrapers table to translators table; all import/export/web translators now belong in this table - adds Scholar.Translate to handle translation issues. eventually, Scholar.Ingester.Document will become part of this interface - adds Scholar_File_Interface (in fileInterface.js) to handle UI for export and eventually import. (David, when you have time, please connect Scholar_File_Interface.exportFile to a button.) - adds an export translator for MODS. all of our metadata, but not our hierarchy (projects, etc.) translates directly and unambiguously into valid MODS. eventually, we can use RDF or another format to handle hierarchy. - adds utilities.getVersion() and utilities.inArray() for simplified scraper coding - fixes minor interface issues with the nifty chrome scraping status window	2006-06-29 00:56:50 +00:00
Simon Kornblith	257ed8f69b	closes #68 , figure out way to have scrapers work for gated resources behind proxies. most institutions use EZProxy for their proxy needs (or a more transparent proxy, which we support natively). this implementation is significantly better than the old one, which refused to work after you'd already logged in once, and is also simpler, because it's stateless. it has to observe every HTTP request, but there's no noticeable speed hit. it also still doesn't work when there's a link from one gated site to another gated site, but as far as i can tell, this only happens on the Gale Group site.	2006-06-27 04:08:21 +00:00
Simon Kornblith	4242c62b1b	- Fix redundancy in utilities.js (I accidentally copied and pasted a much larger block of code than i meant to) - Move processDocuments, a function for loading a DOM representation of a document or set of documents, to Scholar.Utilities.HTTP - Add Scholar.Ingester.ingestURL, a simplified function to scrape a URL (closes #33)	2006-06-26 20:02:30 +00:00
Simon Kornblith	4535b220db	Closes #84 , make type icon in toolbar match item about to be scraped. It's not perfect, since to get everything right, we'd need to scrape the page as soon as it appears, but it provides a pretty good indication. Multiple items get the folder icon. If there's a better icon out there, it's pretty straightforward to implement.	2006-06-26 18:05:23 +00:00
Simon Kornblith	04730860a6	Move Scholar.HTTP to Scholar.Utilities.HTTP; create Scholar.Utilities.Ingester.HTTPUtilities to handle proxied URLs for Ingester	2006-06-26 16:18:55 +00:00
Simon Kornblith	7148852955	make generic Scholar.Utilities class and HTTP-dependent Scholar.Utilities.Ingester and Scholar.Utilities.HTTP classes in preparation for import/export filters; split off into separate javascript file	2006-06-26 14:46:57 +00:00
Simon Kornblith	303c6ee68d	closes #41 , get library call number	2006-06-26 01:08:59 +00:00
Simon Kornblith	5e73dcdd2e	- Search results scraping for WorldCat. - Make scraperJavaScript run on reload again, because it makes debugging easier - There's not actually a memory leak in the proxyMonitor code.	2006-06-25 16:13:47 +00:00
Simon Kornblith	9e78d62b13	Better handling of itemTypes, and improved date handling in PubMed scraper.	2006-06-25 05:03:01 +00:00
Simon Kornblith	22eebc6cdf	Addresses #68 , figure out way to have scrapers work for gated resources behind proxies. We can now access pages through an EZProxy. We need to know what alternatives to EZProxy exist in order to support them. Also, fixes some spacing issues in browser.js.	2006-06-25 04:30:43 +00:00
Simon Kornblith	f897564f0e	Temporary fix to get ingested item types right until #66 is implemented	2006-06-24 21:44:36 +00:00
Simon Kornblith	260ce80086	- Search results scraping for TLC. This is the last of the library scrapers. - Minor fixes to ingester utilities.	2006-06-24 15:38:53 +00:00
Dan Stillman	97940c7470	Replaced all instances of "Firefox Scholar" (not counting the repository URL) with "Scholar for Firefox" for now	2006-06-24 09:08:12 +00:00
Simon Kornblith	2a74e88416	- Make generalized function for finding search results case insensitive - Scrape DRA search results	2006-06-23 20:09:48 +00:00
Simon Kornblith	098078627c	- Make events listening for DOMContentLoaded listen for load, because DOMContentLoaded does not seem ready for prime time (hey, it's undocumented, what can you expect) - Make Amazon scraper work with multiple documents - Fix bugs in processDocuments - Make Scholar.Ingester.Utilities.getItemArray() willing to take an array of DOM nodes to search for links, and finally take advantage of the fact that objects have no length	2006-06-23 03:02:30 +00:00
Simon Kornblith	470f7c463f	The Voyager scraper now actually works on the search results page.	2006-06-22 20:50:57 +00:00
Simon Kornblith	3890e5f122	- Made ingester automatically create hidden browser objects, given a window object. This should make things much easier for both David and me. - Multiple item detection code is now a part of the scraperJavaScript, rather than the scrapeDetectCode, and code to choose which items to add is part of Scholar.Ingester.Utilities, accessible from inside scrapers. The alternative approach would result in one request (or, in the case of JSTOR, three requests) per new item, while in some cases (e.g. Voyager) only one request is necessary to get all of the items.	2006-06-22 15:50:46 +00:00
Simon Kornblith	ca3a0e6e5d	Beginnings of search result scraping (does not yet actually do the scraping, but does present the menu)	2006-06-22 02:43:40 +00:00
Simon Kornblith	6d1e447154	- Remove load eventListener after it has been called once - Capture editors from Google Books	2006-06-21 15:18:18 +00:00
Simon Kornblith	7d3deb5b9f	- Make Scholar.Ingester.Utilities.loadDocument() attach an event handler to load rather than DOMContentLoaded to resolve an issue with the Ex Libris/Aleph scraper (VCU) - When possible, corporate creators/contributors are categorized with their own RDF types (prefixDummy + "corporateCreator/corporateContributor) - Remove extraneous debug code in extensions	2006-06-21 01:41:07 +00:00
Simon Kornblith	09d79d6dd7	Fix overly optimistic JSTOR scraper	2006-06-20 17:06:41 +00:00
Simon Kornblith	20369f41b3	- Move commonly used scraper functions to ingester.js, rather than re-defining them in each scraper. This breaks Piggy Bank compatibility in our scrapers, but we will still be able to export our scrapers in a Piggy Bank compatible form. - Better handling of scraper RDF to item mapping. - Improved date handling. All scrapers now return ISO-style dates when possible.	2006-06-18 19:04:32 +00:00
Simon Kornblith	3d881eec13	- Make scrapers return standard ISO-style YYYY-MM-DD dates. Still need to work on journal article scrapers. - Ingester lets callback function save items, rather than saving them itself. - Better handling of multiple items in API, although no scrapers currently implement this.	2006-06-17 21:21:15 +00:00
Simon Kornblith	076ee0fad2	Add PubMed scraper, fix a few other small bugs	2006-06-08 01:26:40 +00:00
Simon Kornblith	152c9bf9e7	- Small changes to MARC record support - Implemented loadDocument API, for loading and parsing the DOMs of HTML documents in the background - Added scraper code to SVN repository (now includes 12 scrapers, see Writeboard for details) To update to the latest versions of all scrapers, ensure you have an up-to-date version of sqlite3, then run: sqlite3 ~/Library/Application\ Support/Firefox/Profiles/profileName/scholar.sqlite < scrapers.sql	2006-06-06 18:25:45 +00:00
Simon Kornblith	85d8153024	Add library, hooks for scraping MARC records.	2006-06-03 22:26:01 +00:00
Simon Kornblith	93652a137c	Fix issues with asynchronous scraping and XMLHttpRequest	2006-06-02 23:53:42 +00:00
Simon Kornblith	bb57e6ba7d	Provide visual feedback for scraping	2006-06-02 18:22:34 +00:00
Simon Kornblith	639a006efb	XPCOM-ize ingester, fix swapped first and last name in ingested info, stop ingesting pages field (this should be for pages of the source used, not the total number of pages, right?)	2006-06-02 03:19:12 +00:00