relational database

This commit is contained in:
Mehmet D. Akın
2007-03-21 10:20:17 +00:00
parent 892bdce9ac
commit f4d50eb3ae
+34
View File
@@ -82,6 +82,40 @@ and repository name as key and add that dictionary to db. For example if I am ad
packages db, I am actually adding, a Dictionary with ["repotype", foo] to db.
This is cheap, bloated and slow. We must have different repository databases for packages.
* We should really consider using a simple and fast relational database (sqlite) for database operations. Here
is the rationale and some issues.
- Current implementation uses berkeleydb as a storage of dictionaries, with keys and Pickled Python objects.
This approach simplifies coding requirements but leads to some problems.
1. Some operations cannot be done effectively. Searching can only be done by iterating all objects and searching
one by one. component, reverse dependencies etc are holding lists of package names that are updated on every
package operation, Multiple repository problem is
2.Performance: although fast, pickling and unpickling whole objects is still a time consuming process.
Pickled Objects contains object metadata and are larger then the data they hold, so reading, writing pickled
Objects to db is slower than reading and writing only necessary fields. large objects (Order of kilobytes)
causes berkeleydb data buckets to overflow and caches works less efficently.
4.Disk usage: Using pickled objects requires more memory and disk space. Total size of a current pisi database
is over 100Mb, if we modify files.bdb to only hold filenames and packages, size is still around 50MB.
Rational Databases offers number of advantages and disadvantages:
Advantages:
+ Relational databases offers much more flexible operations on stored data. Searching, grouping, sorting
can be done by database engine efficently.
+ There will be only a single database file and much smaller database size compared to berkeleydb files
holding pickled python objects.
+ Since database no longer holds pickled python objects, it is possible to read and modify the database
by other programs written in different languages (Java, C++ etc.)
+ SQL lite is generally very fast, probably beating "Berkeleydb with pickled objects" by the order of 4-5 on
queries.
Disadvantages
- Using a relational database will require more coding when adding and retrieving contents of objects.
- Introducing SQL language makes code more complex and harder to maintain.
- Designing a good schema, identifying required indexes etc. takes time
- Some developers are superstitious about sql and relational databases ;)
==> class attributes
* In some classes there are some attributes assigned but never
used. (see remove_unused_attributes.patch)